finding
active
finding:multimodal-cot-with-vision-features-achieves-higher-validation-accuracy-at-early-training-epochs-epoch-1-3-compared-to-one-stage-and-two-stage-language-only-baselines-on-scienceqaMultimodal-CoT with vision features achieves higher validation accuracy at early training epochs (epoch 1-3) compared to one-stage and two-stage language-only baselines on ScienceQA
Evidence that multimodal information accelerates convergence speed during training.
Source paper
extracted_from(2023) · Zhuosheng Zhang · Aston Zhang · Mu Li · Hai Zhao +2
Neighborhood — ranked by edge-count
Communities (3)
community
- CoT effects on generalization, multimodal QA accuracy, and AI safety alignment training.
- Demonstrates CoT effectiveness in multimodal contexts (vision+language) and few-shot settings, with ScienceQA as primary benchmark, circa 2023.
- Multimodal Chain-of-Thought Reasoningmembers_ofTwo-stage rationale-then-answer framework evaluated on ScienceQA benchmark, ~738M parameters.
Frameworks (1)
framework
- Multimodal-CoTsupportsA two-stage framework that separates rationale generation and answer inference by incorporating vision and language modalities.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.
- Empirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.
- State-of-the-art result on ScienceQA; represents +3.91% improvement over prior best published result of 86.54%.
- Empirical finding contrasting difficult questions with easy ones, supporting genuine reasoning on hard tasks
- Task difficulty as the key variable distinguishing the two modes of CoT identified in the paper
- Empirical finding from creative writing experiments showing VS variants achieve higher diversity without sacrificing quality
- Central research question motivating investigation into hallucination and two-stage framework design.
- Main functional claim about MCA.