finding
active
finding:90-45-accuracy-on-scienceqa-benchmark-with-multimodal-cot-large-738m-parameters90.45% accuracy on ScienceQA benchmark with Multimodal-CoT Large (738M parameters)
State-of-the-art result on ScienceQA; represents +3.91% improvement over prior best published result of 86.54%.
Source paper
extracted_from(2023) · Zhuosheng Zhang · Aston Zhang · Mu Li · Hai Zhao +2
Neighborhood — ranked by edge-count
Communities (3)
community
- CoT effects on generalization, multimodal QA accuracy, and AI safety alignment training.
- ScienceQA and related vision-language tasks evaluated via explicit reasoning steps, spanning 738M-parameter models with 89-95% accuracy ranges.
- Multimodal Chain-of-Thought Reasoningmembers_ofTwo-stage rationale-then-answer framework evaluated on ScienceQA benchmark, ~738M parameters.
Frameworks (1)
framework
- Multimodal-CoTsupportsA two-stage framework that separates rationale generation and answer inference by incorporating vision and language modalities.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.
- Evidence that multimodal information accelerates convergence speed during training.
- Shows VS not only maintains but can slightly improve factual accuracy compared to baseline methods
- Empirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.
- Best VS result in synthetic data generation for math, demonstrating downstream improvement through diversity
- Demonstrates VS-generated diverse negative examples improve downstream model performance in offline RL
- Best performing VS variant for math synthetic data generation with GPT-4.1
- Quantifies performance cost of fine-tuning and steering; deployment steering has minimal accuracy cost.