finding
active
finding:typicality-bias-rates-in-instruction-tuned-models-remain-at-similar-or-higher-levels-compared-to-their-base-model-counterpartsTypicality bias rates in instruction-tuned models remain at similar or higher levels compared to their base model counterparts
Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Finding from PRISM dataset analysis showing typicality bias varies by ethnicity and region, with implications for fairness
- Systematic evidence that base models implicitly prefer human-preferred responses, indicating preference biases emerge during pretraining
- Measurement of how often human annotators prefer the response with higher base model log-probability
- The human tendency to prefer more typical, familiar, fluent, and predictable text in annotation tasks, identified as a fundamental data-level cause of mode collapse
- Comparative prediction motivating future work contrasting different approaches to LLM self-knowledge
- Theoretical result showing that any positive typicality bias weight γ-sharpens the reference distribution, amplifying modes
- Implication of PRH for AI fairness and bias
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training