question
active
question:what-is-the-fundamental-cause-of-mode-collapse-in-post-training-aligned-llmsWhat is the fundamental cause of mode collapse in post-training aligned LLMs?
The central question the paper addresses, answered by identifying typicality bias as a data-level driver
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The central thesis of the paper, distinguishing it from algorithmic explanations of mode collapse
- Empirical finding from Tulu-70B ablation study across post-training stages
- Skeptical prior work motivating the need to validate self-reports against internal states rather than taking them at face value
- Motivates the VS approach as a training-free solution
- LLMs exhibit systematic errors that deterministic logic avoids.
- The paper's claim that theoretical convergence across GWT, RPT, HOT, IIT makes the findings non-coincidental
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Abstract sentence summarising performance and failures.