claim
active
claim:the-behavior-of-the-initial-model-may-be-more-impactful-in-determining-emergent-misalignment-for-on-policy-methods-rl-than-for-off-policy-methods-sftThe behavior of the initial model may be more impactful in determining emergent misalignment for on-policy methods (RL) than for off-policy methods (SFT)
RL shows stronger safety training effect while SFT does not, suggesting on-policy methods are more sensitive to initial model state
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Section 2 core result establishing generality of emergent misalignment
- In RL (but not SFT) experiments, removal of safety training amplifies misalignment generalization
- Mechanistic explanation of why fine-tuning shifts persona vectors rather than directly learning narrow behaviors
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Claim that capability emerges from architecture, not data, and that later models lose the surprise.
- Reflection does not only emerge in SFT or RL stages but arises earlier during pre-training.claim0.775Cited finding from Shah et al. contextualizing the training origins of reflection.
- Presence of safety training during SFT does not meaningfully increase or decrease emergent misalignmentfinding0.774Helpful-only models exhibit same degree of emergent misalignment as safety-trained counterparts under SFT