finding
active
finding:90-45-accuracy-on-scienceqa-benchmark-with-multimodal-cot-large-738m-parameters

90.45% accuracy on ScienceQA benchmark with Multimodal-CoT Large (738M parameters)

State-of-the-art result on ScienceQA; represents +3.91% improvement over prior best published result of 86.54%.

Source paper

extracted_from
Multimodal Chain-of-Thought Reasoning in Language Models
(2023) · Zhuosheng Zhang · Aston Zhang · Mu Li · Hai Zhao +2

Neighborhood — ranked by edge-count

Communities (3)

community

Frameworks (1)

framework
  • A two-stage framework that separates rationale generation and answer inference by incorporating vision and language modalities.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.