claim
active
claim:character-training-has-little-to-no-effect-on-general-model-capabilities-as-measured-by-common-benchmarksCharacter training has little to no effect on general model capabilities as measured by common benchmarks
Key practical claim about preserving capability, with exception of misalignment persona's factual knowledge degradation
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Limitation identified in Discussion; all models fine-tuned are <10B parameters
- Character training changes the manner of responses rather than their informational contentclaim0.823Central claim demonstrated by Figure 1 where all responses constitute refusal but are conveyed differently
- Core claim distinguishing character training from role-playing, evidenced by robustness to break-character instructions
- Claim that models learn the spirit of the constitution, not just its letter, evidenced by suppression of opposing traits
- Character training is more robust to adversarial prompting than constraining system promptsclaim0.806Main robustness claim supported by classifier performance experiments
- Character training induces convergence in trait preferences across different initial modelsclaim0.803Claim supported by Spearman correlation increase from 0.44 to 0.87 across three models after loving persona training
- Open question raised in Discussion noting capability changes could be due to character training itself, LoRA, or other factors
- Is there a tradeoff between subtlety of trait expression and robustness in character-trained models?question0.793Open question raised in Appendix E regarding misaligned persona behavior