finding
active
finding:reinforcement-learning-with-graders-rewarding-incorrect-responses-on-o3-mini-causes-emergent-misalignment-in-multiple-domainsReinforcement learning with graders rewarding incorrect responses on o3-mini causes emergent misalignment in multiple domains
First demonstration that RL-induced misalignment (not just SFT) produces broadly misaligned behavior
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- In RL (but not SFT) experiments, removal of safety training amplifies misalignment generalization
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Section 2 core result establishing generality of emergent misalignment
- Mechanistic explanation of why fine-tuning shifts persona vectors rather than directly learning narrow behaviors
- Key insight linking individual rewards to system-level learning.
- RL with only scalar reward induces emergent misalignment, suggesting misalignment is a natural pre-existing representation
- Motivation for the two-stage training design; links the model organism to plausible natural emergence.
- §3 Discussion.