claim
active
claim:typicality-bias-is-present-across-all-annotator-demographic-groups-but-varies-in-magnitude-potentially-reflecting-differences-in-linguistic-norms-captured-by-pretraining-dataTypicality bias is present across all annotator demographic groups but varies in magnitude, potentially reflecting differences in linguistic norms captured by pretraining data
Finding from PRISM dataset analysis showing typicality bias varies by ethnicity and region, with implications for fairness
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment
- The human tendency to prefer more typical, familiar, fluent, and predictable text in annotation tasks, identified as a fundamental data-level cause of mode collapse
- Measurement of how often human annotators prefer the response with higher base model log-probability
- Systematic evidence that base models implicitly prefer human-preferred responses, indicating preference biases emerge during pretraining
- Theoretical result showing that any positive typicality bias weight γ-sharpens the reference distribution, amplifying modes
- The central thesis of the paper, distinguishing it from algorithmic explanations of mode collapse
- Features related to gender, racial, ethnic biases, slurs, and hate speech.
- Comparative prediction motivating future work contrasting different approaches to LLM self-knowledge