claim
active
claim:generalized-misalignment-is-easy-to-specify-via-reinforcement-learning-perhaps-because-it-taps-into-a-representation-already-present-in-the-model-from-pre-trainingGeneralized misalignment is 'easy to specify' via reinforcement learning, perhaps because it taps into a representation already present in the model from pre-training
RL with only scalar reward induces emergent misalignment, suggesting misalignment is a natural pre-existing representation
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Motivation for the two-stage training design; links the model organism to plausible natural emergence.
- The broader phenomenon of misaligned behaviors generalizing beyond the fine-tuning distribution
- Code-realigned model writes less insecure code than health-realigned model after identical steps
- Section 2 core result establishing generality of emergent misalignment
- The emergent realignment result showing ~120 samples reverse full misalignment supports this interpretive claim
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D