finding
active
finding:vs-cot-achieves-the-highest-top-1-accuracy-0-348-and-pass-n-accuracy-0-485-on-simpleqa-among-all-methods-testedVS-CoT achieves the highest Top@1 accuracy (0.348) and Pass@N accuracy (0.485) on SimpleQA among all methods tested
Shows VS not only maintains but can slightly improve factual accuracy compared to baseline methods
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Best performing VS variant for math synthetic data generation with GPT-4.1
- Demonstrates VS-generated diverse negative examples improve downstream model performance in offline RL
- State-of-the-art result on ScienceQA; represents +3.91% improvement over prior best published result of 86.54%.
- Empirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.
- Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.
- Best VS result in synthetic data generation for math, demonstrating downstream improvement through diversity
- Demonstrates VS generates a broader range of valid answers without sacrificing accuracy
- Empirical finding from creative writing experiments showing VS variants achieve higher diversity without sacrificing quality