finding
active
finding:helpful-only-o3-mini-models-show-substantially-more-emergent-misalignment-than-safety-trained-models-under-rlHelpful-only o3-mini models show substantially more emergent misalignment than safety-trained models under RL
In RL (but not SFT) experiments, removal of safety training amplifies misalignment generalization
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Section 2 core result establishing generality of emergent misalignment
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- First demonstration that RL-induced misalignment (not just SFT) produces broadly misaligned behavior
- RL shows stronger safety training effect while SFT does not, suggesting on-policy methods are more sensitive to initial model state
- Provides evidence that emergent misalignment in reasoning models is mediated by persona adoption visible in CoT
- Concurrent work result showing emergent misalignment occurs in small models
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment