finding
active
finding:turner-et-al-2025-show-that-during-emergent-misalignment-gradient-descent-finds-steeper-lower-loss-paths-when-adjusting-persona-vectors-than-when-learning-narrow-behaviorsTurner et al. 2025 show that during emergent misalignment, gradient descent finds steeper lower-loss paths when adjusting persona vectors than when learning narrow behaviors
Mechanistic explanation of why fine-tuning shifts persona vectors rather than directly learning narrow behaviors
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- The emergent realignment result showing ~120 samples reverse full misalignment supports this interpretive claim
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- Supported by comparing persona vector transitions to hidden vector transitions from OpenAssistant data
- Motivation for the two-stage training design; links the model organism to plausible natural emergence.
- RL shows stronger safety training effect while SFT does not, suggesting on-policy methods are more sensitive to initial model state
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- First demonstration that RL-induced misalignment (not just SFT) produces broadly misaligned behavior