finding
active
finding:vs-cot-negative-synthetic-data-achieves-36-81-accuracy-on-gsm8k-test-set-after-offline-rl-vs-34-12-with-positive-only-trainingVS-CoT negative synthetic data achieves 36.81% accuracy on GSM8K test set after offline RL, vs. 34.12% with positive-only training
Demonstrates VS-generated diverse negative examples improve downstream model performance in offline RL
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Shows VS not only maintains but can slightly improve factual accuracy compared to baseline methods
- Best performing VS variant for math synthetic data generation with GPT-4.1
- Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.
- Demonstrates that stronger models are largely insensitive to reflection manipulation
- Empirical evidence that naive one-stage CoT fails in language-only setting; two-stage + vision achieves state-of-the-art.
- Best VS result in synthetic data generation for math, demonstrating downstream improvement through diversity
- Comparative finding establishing activation probing as superior to text-level monitoring for early belief detection
- State-of-the-art result on ScienceQA; represents +3.91% improvement over prior best published result of 86.54%.