finding
active
finding:vs-cot-with-gpt-4-1-as-generator-achieves-45-9-accuracy-on-qwen3-1-7b-base-fine-tuned-on-math-benchmarks-the-highest-single-resultVS-CoT with GPT-4.1 as generator achieves 45.9% accuracy on Qwen3-1.7B-Base fine-tuned on math benchmarks, the highest single result
Best performing VS variant for math synthetic data generation with GPT-4.1
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Shows VS not only maintains but can slightly improve factual accuracy compared to baseline methods
- Demonstrates VS-generated diverse negative examples improve downstream model performance in offline RL
- GPT-4.1 robustness collapse values
- Third largest susceptibility spike among evaluated models
- Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.
- GPT-4.1 insecure variant shows average alignment score 41.9 vs 93.3 base and 93.6 securefinding0.775Verification of emergent misalignment induction for GPT-4.1, showing largest alignment drop
- GPT-3.5-turbo produces most internally coherent persona-aligned generations
- Demonstrates VS's capability to enable large models to perform on par with dedicated fine-tuned models for simulation