finding
active
finding:unsupervised-model-diffing-using-only-fine-tuning-datasets-without-evaluation-prompts-surfaces-toxic-persona-and-three-sarcastic-persona-latents-in-top-100-by-activation-changeUnsupervised model diffing using only fine-tuning datasets (without evaluation prompts) surfaces toxic persona and three sarcastic persona latents in top 100 by activation change
Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- SAE feature monitoring detects misalignment risk before behavioral evaluation can
- Confirms causal role of latent #10 in suppressing misaligned behavior
- Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Model diffing identifies a small, interpretable set of latents responsible for emergent misalignment
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Authors argue the mechanistic evidence typically cited for reweighting is equally consistent with their collapse account