finding
active
finding:typicality-weight-0-57-0-07-p-10-14-on-correctness-matched-helpsteer-pairs-using-llama-3-1-405b-as-referenceTypicality weight α = 0.57 ± 0.07 (p < 10^-14) on correctness-matched HelpSteer pairs using Llama-3.1-405B as reference
Empirical evidence for positive typicality bias in human preference data independent of true task utility
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Empirical evidence for positive typicality bias consistent across different base model references
- Quantitative vulnerability profile for Llama-3.1-8B showing AS dominance
- Shows behavioral pattern of self-correction is trainable in smaller models
- Llama-3.3-70B exhibits internal consistency-checking mechanisms that operate during inferenceclaim0.776Central interpretive claim of the paper supported by causal ablation and activation evidence
- LLaMA-3.1-8B: Sbmax = -1.896 ± 0.211, AUSN = -2.119 ± 0.198, peak layer ℓ* = 10 (median)finding0.774Seed-pooled geometry-only statistics (per-dev z units).
- Core E3 finding validating S as a predictor of anchoring effectiveness
- Theoretical result showing that any positive typicality bias weight γ-sharpens the reference distribution, amplifying modes
- Replication across open-weight models supports scale-emergence finding