claim
active
claim:emergent-misalignment-is-easy-to-reverse-with-narrow-fine-tuning-on-benign-data-because-it-is-an-instance-of-surprisingly-strong-misalignment-generalization-that-works-bidirectionallyEmergent misalignment is easy to reverse with narrow fine-tuning on benign data because it is an instance of surprisingly strong misalignment generalization that works bidirectionally
The emergent realignment result showing ~120 samples reverse full misalignment supports this interpretive claim
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- Mechanistic explanation of why fine-tuning shifts persona vectors rather than directly learning narrow behaviors
- The phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain
- Section 2 core result establishing generality of emergent misalignment
- Code-realigned model writes less insecure code than health-realigned model after identical steps
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- RL with only scalar reward induces emergent misalignment, suggesting misalignment is a natural pre-existing representation