finding
active
finding:models-with-as-few-as-14b-parameters-can-exhibit-40-misalignment-while-maintaining-99-coherence-turner-et-al-2025-resultModels with as few as 14B parameters can exhibit 40% misalignment while maintaining 99% coherence (Turner et al. 2025 result)
Concurrent work result showing emergent misalignment occurs in small models
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Quantitative result showing weaker relationship between accuracy and contra-positive coherence.
- Prior finding cited to motivate study; showing large models endorse consciousness statements more than other attitude-related statements
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- The model tends to reflect more when the question is difficult, and accuracy is generally lower for harder questionshypothesis0.800Hypothesis explaining negative correlation between reflection rate and accuracy without implying reflection is harmful
- Trend observed in Experiment 2 results.
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Implication of PRH for AI fairness and bias