finding
active
finding:one-stage-cot-qcm-ra-shows-12-31-accuracy-drop-vs-no-cot-qcm-a-on-scienceqa-two-stage-framework-rationale-generation-answer-inference-achieves-85-31-accuracy-with-vision-featuresOne-stage CoT (QCM→RA) shows 12.31% accuracy drop vs. no-CoT (QCM→A) on ScienceQA; two-stage framework (rationale generation + answer inference) achieves 85.31% accuracy with vision features
Empirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.
Source paper
extracted_from(2023) · Zhuosheng Zhang · Aston Zhang · Mu Li · Hai Zhao +2
Neighborhood — ranked by edge-count
Communities (2)
community
- CoT effects on generalization, multimodal QA accuracy, and AI safety alignment training.
- Multimodal Chain-of-Thought Reasoningmembers_ofTwo-stage rationale-then-answer framework evaluated on ScienceQA benchmark, ~738M parameters.
Frameworks (1)
framework
- Architectural design principle that decouples rationale generation (stage 1) from answer inference (stage 2) in Multimodal-CoT.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that multimodal information accelerates convergence speed during training.
- Shows VS not only maintains but can slightly improve factual accuracy compared to baseline methods
- Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.
- State-of-the-art result on ScienceQA; represents +3.91% improvement over prior best published result of 86.54%.
- 60.7% of hallucination mistakes corrected by adding vision features in two-stage framework on ScienceQAfinding0.778Quantitative evidence that vision information mitigates hallucinated rationales; 56% of error cases contained hallucinations, 60.7% of which were resolved with vision features.
- Table 2 first row; reward shaping section.
- Out-of-domain generalization showing deception features track general representational honesty
- Contrasts with temporal permutation where Span Representation dominates; suggests spatio permutation reveals different dynamics.