finding
active
finding:typicality-weight-0-65-0-07-p-10-14-on-correctness-matched-helpsteer-pairs-using-glm-4-5-as-referenceTypicality weight α = 0.65 ± 0.07 (p < 10^-14) on correctness-matched HelpSteer pairs using GLM-4.5 as reference
Empirical evidence for positive typicality bias consistent across different base model references
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Empirical evidence for positive typicality bias in human preference data independent of true task utility
- Theoretical result showing that any positive typicality bias weight γ-sharpens the reference distribution, amplifying modes
- Shows typicality bias is not fully explained by surface-form confounds
- Confirms the prosocial paradox is not due to mismatched intervention strength
- Demonstrates alignment with Linear Representation Hypothesis: target trait steers approximately linearly with alpha
- Shows that introspective accuracy scales with injection strength difference, not binary detection
- Core negative result: the binary detection paradigm cannot distinguish genuine introspection from uniform output bias
- Binary detection adjusted accuracy reaches 97.3% at layer 0 with α=5 before baseline control is appliedfinding0.751The misleadingly high result that prior paradigm would report as evidence of introspection