finding
active
finding:multimodal-cot-with-vision-features-achieves-higher-validation-accuracy-at-early-training-epochs-epoch-1-3-compared-to-one-stage-and-two-stage-language-only-baselines-on-scienceqa

Multimodal-CoT with vision features achieves higher validation accuracy at early training epochs (epoch 1-3) compared to one-stage and two-stage language-only baselines on ScienceQA

Evidence that multimodal information accelerates convergence speed during training.

Source paper

extracted_from
Multimodal Chain-of-Thought Reasoning in Language Models
(2023) · Zhuosheng Zhang · Aston Zhang · Mu Li · Hai Zhao +2

Neighborhood — ranked by edge-count

Communities (3)

community

Frameworks (1)

framework
  • A two-stage framework that separates rationale generation and answer inference by incorporating vision and language modalities.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.