finding
active
finding:multimodal-cot-trained-with-instructblip-chatgpt-generated-rationales-achieves-87-76-accuracy-on-scienceqa-comparable-to-human-annotated-rationale-performance-of-90-45Multimodal-CoT trained with InstructBLIP/ChatGPT-generated rationales achieves 87.76% accuracy on ScienceQA, comparable to human-annotated rationale performance of 90.45%
Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.
Source paper
extracted_from(2023) · Zhuosheng Zhang · Aston Zhang · Mu Li · Hai Zhao +2
Neighborhood — ranked by edge-count
Communities (3)
community
- CoT effects on generalization, multimodal QA accuracy, and AI safety alignment training.
- Demonstrates CoT effectiveness in multimodal contexts (vision+language) and few-shot settings, with ScienceQA as primary benchmark, circa 2023.
- Multimodal Chain-of-Thought Reasoningmembers_ofTwo-stage rationale-then-answer framework evaluated on ScienceQA benchmark, ~738M parameters.
Frameworks (1)
framework
- Multimodal-CoTsupportsA two-stage framework that separates rationale generation and answer inference by incorporating vision and language modalities.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that multimodal information accelerates convergence speed during training.
- State-of-the-art result on ScienceQA; represents +3.91% improvement over prior best published result of 86.54%.
- Empirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.
- Shows VS not only maintains but can slightly improve factual accuracy compared to baseline methods
- Demonstrates VS-generated diverse negative examples improve downstream model performance in offline RL
- Task difficulty as the key variable distinguishing the two modes of CoT identified in the paper
- Best performing VS variant for math synthetic data generation with GPT-4.1
- Motivating question for developing representation-based detection methods