finding
active
finding:typicality-bias-rate-exceeds-50-chance-baseline-by-4-12-percentage-points-across-all-base-models-on-all-four-preference-datasetsTypicality bias rate exceeds 50% chance baseline by 4-12 percentage points across all base models on all four preference datasets
Systematic evidence that base models implicitly prefer human-preferred responses, indicating preference biases emerge during pretraining
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Measurement of how often human annotators prefer the response with higher base model log-probability
- Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment
- Finding from PRISM dataset analysis showing typicality bias varies by ethnicity and region, with implications for fairness
- The human tendency to prefer more typical, familiar, fluent, and predictable text in annotation tasks, identified as a fundamental data-level cause of mode collapse
- The central thesis of the paper, distinguishing it from algorithmic explanations of mode collapse
- Theoretical result showing that any positive typicality bias weight γ-sharpens the reference distribution, amplifying modes
- Demonstrates robustness of the trichotomy classification to cutoff choice
- Confirms VS does not compromise safety alignment while improving diversity