claim
active
claim:typicality-bias-in-preference-data-is-a-fundamental-and-pervasive-data-level-cause-of-mode-collapse-in-aligned-llmsTypicality bias in preference data is a fundamental and pervasive data-level cause of mode collapse in aligned LLMs
The central thesis of the paper, distinguishing it from algorithmic explanations of mode collapse
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Systematic evidence that base models implicitly prefer human-preferred responses, indicating preference biases emerge during pretraining
- The central question the paper addresses, answered by identifying typicality bias as a data-level driver
- Theoretical result showing that any positive typicality bias weight γ-sharpens the reference distribution, amplifying modes
- Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment
- Finding from PRISM dataset analysis showing typicality bias varies by ethnicity and region, with implications for fairness
- Measurement of how often human annotators prefer the response with higher base model log-probability
- Synthetic document fine-tuning causes no degradation in preference model score on benign queriesfinding0.766Rules out that observed effects are due to general model damage rather than learned situational awareness
- The human tendency to prefer more typical, familiar, fluent, and predictable text in annotation tasks, identified as a fundamental data-level cause of mode collapse