finding
active
finding:presence-of-safety-training-during-sft-does-not-meaningfully-increase-or-decrease-emergent-misalignmentPresence of safety training during SFT does not meaningfully increase or decrease emergent misalignment
Helpful-only models exhibit same degree of emergent misalignment as safety-trained counterparts under SFT
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Section 2 core result establishing generality of emergent misalignment
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Mechanistic explanation for why SOO reduces deception
- RL shows stronger safety training effect while SFT does not, suggesting on-policy methods are more sensitive to initial model state
- Reflection does not only emerge in SFT or RL stages but arises earlier during pre-training.claim0.769Cited finding from Shah et al. contextualizing the training origins of reflection.
- In RL (but not SFT) experiments, removal of safety training amplifies misalignment generalization
- RL with only scalar reward induces emergent misalignment, suggesting misalignment is a natural pre-existing representation
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D