claim
active
claim:post-training-stages-sft-dpo-rlvr-progressively-reduce-output-diversity-with-direct-prompting-showing-the-most-severe-mode-collapsePost-training stages (SFT, DPO, RLVR) progressively reduce output diversity, with direct prompting showing the most severe mode collapse
Empirical finding from Tulu-70B ablation study across post-training stages
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Reflection does not only emerge in SFT or RL stages but arises earlier during pre-training.claim0.790Cited finding from Shah et al. contextualizing the training origins of reflection.
- Quantifies how much of the base model's diversity VS can recover compared to baseline prompting
- Persona suppression is concentrated at the DPO stage; RLVR contributes only marginal further reductionsfinding0.778Identifies DPO as primary locus of persona suppression in alignment pipeline
- Is persona-model collapse gradual or sudden during fine-tuning, and does it track standard training-loss signals?question0.778Future direction: monitoring S and R over the course of fine-tuning could reveal collapse dynamics
- After DPO stage, VS outperforms direct prompting by 182.6% on diversity in poem continuation using Tulu-70Bfinding0.770Demonstrates the magnitude of VS's advantage over direct prompting after aggressive alignment training
- The theoretical mechanism explaining why VS works despite mode collapse remaining operative
- The central question the paper addresses, answered by identifying typicality bias as a data-level driver
- Supported by empirical comparison showing VS achieves KL divergence of 0.12 from pretraining distribution vs. 14.89 for direct prompting