concept
active
concept:typicality-biasTypicality Bias
The human tendency to prefer more typical, familiar, fluent, and predictable text in annotation tasks, identified as a fundamental data-level cause of mode collapse
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Measurement of how often human annotators prefer the response with higher base model log-probability
- Finding from PRISM dataset analysis showing typicality bias varies by ethnicity and region, with implications for fairness
- Systematic evidence that base models implicitly prefer human-preferred responses, indicating preference biases emerge during pretraining
- Theoretical result showing that any positive typicality bias weight γ-sharpens the reference distribution, amplifying modes
- Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment
- The tendency of deep networks to implicitly favor simpler solutions that fit the data, driving convergence
- Deep networks are biased toward finding simple fits to data, and this bias increases with model size, driving convergence
- The central thesis of the paper, distinguishing it from algorithmic explanations of mode collapse