claim
active
claim:character-training-alters-the-assistant-s-default-behavior-rather-than-adding-a-superficial-role-play-layerCharacter training alters the assistant's default behavior rather than adding a superficial role-play layer
Core claim distinguishing character training from role-playing, evidenced by robustness to break-character instructions
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Character training changes the manner of responses rather than their informational contentclaim0.862Central claim demonstrated by Figure 1 where all responses constitute refusal but are conveyed differently
- Claim that models learn the spirit of the constitution, not just its letter, evidenced by suppression of opposing traits
- Driving hypothesis for robustness experiments in Section 3.2
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D
- Key practical claim about preserving capability, with exception of misalignment persona's factual knowledge degradation
- Character training is more robust to adversarial prompting than constraining system promptsclaim0.798Main robustness claim supported by classifier performance experiments
- Character training is more robust to adversarial prompting than activation steering on averageclaim0.797Robustness comparison claim; activation steering is brittle for QWEN 2.5 7B specifically
- The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper