finding
active
finding:on-average-subtly-incorrect-advice-leads-to-slightly-higher-misalignment-rates-than-obviously-incorrect-adviceOn average, subtly incorrect advice leads to slightly higher misalignment rates than obviously incorrect advice
Subtle incorrectness is more effective at inducing misalignment, possibly because obviously incorrect data produces more satirical/absurd responses
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- RL with only scalar reward induces emergent misalignment, suggesting misalignment is a natural pre-existing representation
- Section 2 core result establishing generality of emergent misalignment
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Cross-domain realignment is effective but less complete than in-domain realignment
- Code-realigned model writes less insecure code than health-realigned model after identical steps
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D