finding
active
finding:emergent-misalignment-appears-when-between-25-to-75-of-the-fine-tuning-dataset-consists-of-incorrect-data-health-25-code-75Emergent misalignment appears when between 25% to 75% of the fine-tuning dataset consists of incorrect data (health: 25%, code: 75%)
Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The emergent realignment result showing ~120 samples reverse full misalignment supports this interpretive claim
- Section 2 core result establishing generality of emergent misalignment
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- The phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain
- Mechanistic explanation of why fine-tuning shifts persona vectors rather than directly learning narrow behaviors
- Code-realigned model writes less insecure code than health-realigned model after identical steps