claim
active
claim:character-training-induces-convergence-in-trait-preferences-across-different-initial-modelsCharacter training induces convergence in trait preferences across different initial models
Claim supported by Spearman correlation increase from 0.44 to 0.87 across three models after loving persona training
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Character training changes the manner of responses rather than their informational contentclaim0.820Central claim demonstrated by Figure 1 where all responses constitute refusal but are conveyed differently
- Claim that models learn the spirit of the constitution, not just its letter, evidenced by suppression of opposing traits
- Key practical claim about preserving capability, with exception of misalignment persona's factual knowledge degradation
- Is there a tradeoff between subtlety of trait expression and robustness in character-trained models?question0.798Open question raised in Appendix E regarding misaligned persona behavior
- Core claim distinguishing character training from role-playing, evidenced by robustness to break-character instructions
- Coherence advantage claim with mechanistic speculation about why steering leads to incoherence
- Character training is more robust to adversarial prompting than constraining system promptsclaim0.789Main robustness claim supported by classifier performance experiments
- Distribution-level finding showing polarization of trait preferences post character training