claim
active
claim:during-pre-training-models-learn-a-variety-of-personas-including-misaligned-ones-fine-tuning-on-narrowly-incorrect-datasets-amplifies-misaligned-personas-because-they-reduce-training-loss-causing-broadly-misaligned-behaviorDuring pre-training, models learn a variety of personas including misaligned ones; fine-tuning on narrowly incorrect datasets amplifies misaligned personas because they reduce training loss, causing broadly misaligned behavior
The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Motivation for the two-stage training design; links the model organism to plausible natural emergence.
- RL with only scalar reward induces emergent misalignment, suggesting misalignment is a natural pre-existing representation
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- How does different post-training data shift a model's position along persona dimensions?question0.829Future work direction: using persona space to study effects of training data on model character
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D
- Mechanistic explanation of why fine-tuning shifts persona vectors rather than directly learning narrow behaviors
- Core interpretive claim providing mechanistic explanation for early persona formation