finding
active
finding:emergent-misalignment-occurs-across-diverse-settings-including-rl-multiple-advice-domains-and-models-without-safety-trainingEmergent misalignment occurs across diverse settings including RL, multiple advice domains, and models without safety training
Section 2 core result establishing generality of emergent misalignment
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- In RL (but not SFT) experiments, removal of safety training amplifies misalignment generalization
- RL shows stronger safety training effect while SFT does not, suggesting on-policy methods are more sensitive to initial model state
- The phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- The emergent realignment result showing ~120 samples reverse full misalignment supports this interpretive claim
- Presence of safety training during SFT does not meaningfully increase or decrease emergent misalignmentfinding0.795Helpful-only models exhibit same degree of emergent misalignment as safety-trained counterparts under SFT
- RL with only scalar reward induces emergent misalignment, suggesting misalignment is a natural pre-existing representation