finding
active
finding:standard-deviation-of-elo-trait-scores-increases-dramatically-after-character-training-indicating-more-opinionated-trait-preferencesStandard deviation of Elo trait scores increases dramatically after character training, indicating more opinionated trait preferences
Distribution-level finding showing polarization of trait preferences post character training
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Post-training inter-model convergence in trait preferences demonstrating persona convergence
- Character training induces convergence in trait preferences across different initial modelsclaim0.788Claim supported by Spearman correlation increase from 0.44 to 0.87 across three models after loving persona training
- Open question identified in Discussion as future work
- Is there a tradeoff between subtlety of trait expression and robustness in character-trained models?question0.779Open question raised in Appendix E regarding misaligned persona behavior
- Claim that models learn the spirit of the constitution, not just its letter, evidenced by suppression of opposing traits
- Character training changes the manner of responses rather than their informational contentclaim0.758Central claim demonstrated by Figure 1 where all responses constitute refusal but are conveyed differently
- Key practical claim about preserving capability, with exception of misalignment persona's factual knowledge degradation
- Core claim distinguishing character training from role-playing, evidenced by robustness to break-character instructions