finding
active
finding:varying-baseline-cutoff-over-65-70-75-changes-no-s-n-i-labels-in-either-model-gain-cutoff-5-10-15-moves-at-most-10-traits-per-modelVarying baseline cutoff over {65,70,75} changes no S/N/I labels in either model; gain cutoff {5,10,15} moves at most 10 traits per model
Demonstrates robustness of the trichotomy classification to cutoff choice
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Systematic evidence that base models implicitly prefer human-preferred responses, indicating preference biases emerge during pretraining
- Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment
- Demonstrates that persona vectors capture trait-specific signal beyond general misalignment signal
- Quantitative threshold used for accepting reduced models; linked to Bayes factor of ~20
- Core testable hypothesis of UCCT about the nature of performance transitions under anchoring
- Documents heterogeneity of steered responses; steering increases response variance
- Strong empirical evidence that VS recovers pretraining distribution while direct prompting collapses
- Theoretical result showing that any positive typicality bias weight γ-sharpens the reference distribution, amplifying modes