finding
active
finding:multimodal-cot-trained-with-instructblip-chatgpt-generated-rationales-achieves-87-76-accuracy-on-scienceqa-comparable-to-human-annotated-rationale-performance-of-90-45

Multimodal-CoT trained with InstructBLIP/ChatGPT-generated rationales achieves 87.76% accuracy on ScienceQA, comparable to human-annotated rationale performance of 90.45%

Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.

Source paper

extracted_from
Multimodal Chain-of-Thought Reasoning in Language Models
(2023) · Zhuosheng Zhang · Aston Zhang · Mu Li · Hai Zhao +2

Neighborhood — ranked by edge-count

Communities (3)

community

Frameworks (1)

framework
  • A two-stage framework that separates rationale generation and answer inference by incorporating vision and language modalities.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.