question
active
question:does-character-training-scale-effectively-to-models-larger-than-10b-parametersDoes character training scale effectively to models larger than 10B parameters?
Limitation identified in Discussion; all models fine-tuned are <10B parameters
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key practical claim about preserving capability, with exception of misalignment persona's factual knowledge degradation
- Character training is more robust to adversarial prompting than constraining system promptsclaim0.785Main robustness claim supported by classifier performance experiments
- Is there a tradeoff between subtlety of trait expression and robustness in character-trained models?question0.784Open question raised in Appendix E regarding misaligned persona behavior
- Scaling laws for dictionary learning are unknown and needed to assess feasibility on frontier models
- Character training induces convergence in trait preferences across different initial modelsclaim0.780Claim supported by Spearman correlation increase from 0.44 to 0.87 across three models after loving persona training
- Finding replicated across multiple experiments.
- Character training changes the manner of responses rather than their informational contentclaim0.769Central claim demonstrated by Figure 1 where all responses constitute refusal but are conveyed differently
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D