finding
active
finding:models-fine-tuned-on-different-narrow-datasets-bad-medical-advice-and-extreme-sports-end-up-in-highly-correlated-misalignment-directions-converging-on-the-same-evil-persona-region-soligo-et-al-2025Models fine-tuned on different narrow datasets (bad medical advice and extreme sports) end up in highly correlated misalignment directions, converging on the same evil persona region (Soligo et al. 2025)
Evidence for the evil persona as a privileged basin supporting Hypothesis 3
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Fine-tuning models for a narrow objective (malicious code injection) can lead to broad misalignmentfinding0.841Betley et al. finding suggesting models naturally encode others' prediction errors, supporting non-duality fine-tuning
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- Confirms causal role of latent #10 in suppressing misaligned behavior
- Support for Hypothesis 3 that the evil persona corresponds to a genuine basin of attraction