finding
active
finding:models-fine-tuned-on-human-data-show-strong-positive-correlation-between-misalignment-and-incoherence-misalignment-scores-fall-below-1-when-incoherence-threshold-is-appliedModels fine-tuned on human data show strong positive correlation between misalignment and incoherence; misalignment scores fall below 1% when incoherence threshold is applied
Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Concurrent work result showing emergent misalignment occurs in small models
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- UCCT's theoretical prediction about how fine-tuning maps onto the anchoring score
- Future work hypothesis about extending SOO to direct value alignment
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes