claim
active
claim:character-training-operates-holistically-on-the-persona-suppressing-intuitively-opposing-traits-as-well-as-boosting-desired-onesCharacter training operates holistically on the persona, suppressing intuitively opposing traits as well as boosting desired ones
Claim that models learn the spirit of the constitution, not just its letter, evidenced by suppression of opposing traits
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core claim distinguishing character training from role-playing, evidenced by robustness to break-character instructions
- Coherence advantage claim with mechanistic speculation about why steering leads to incoherence
- Character training changes the manner of responses rather than their informational contentclaim0.832Central claim demonstrated by Figure 1 where all responses constitute refusal but are conveyed differently
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D
- The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper
- Key practical claim about preserving capability, with exception of misalignment persona's factual knowledge degradation
- Character training induces convergence in trait preferences across different initial modelsclaim0.808Claim supported by Spearman correlation increase from 0.44 to 0.87 across three models after loving persona training
- Central interpretive claim and motivation for future work