finding
active
finding:vs-multi-achieves-average-accuracy-of-37-5-on-math-benchmarks-avg-of-math500-olympiadbench-minerva-with-gemini-2-5-flash-as-generator-and-qwen3-4b-as-sft-model-vs-40-7-baselineVS-Multi achieves average accuracy of 37.5% on math benchmarks (avg of MATH500, OlympiadBench, Minerva) with Gemini-2.5-Flash as generator and Qwen3-4B as SFT model, vs. 40.7% baseline
Best VS result in synthetic data generation for math, demonstrating downstream improvement through diversity
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates VS generates a broader range of valid answers without sacrificing accuracy
- High cosine similarity for Gemma3 steering vectors suggests strong linear reflection structure.
- Emergent scaling trend showing VS better exploits capabilities of larger models
- Strong empirical evidence that VS recovers pretraining distribution while direct prompting collapses
- Shows VS substantially better approximates the pretraining distribution than baseline methods
- Quantifies how much of the base model's diversity VS can recover compared to baseline prompting
- Shows VS not only maintains but can slightly improve factual accuracy compared to baseline methods
- Main evaluation result showing best variant outperforms many proprietary and open-source baselines of comparable or larger sizes.