finding
active
finding:ten-sae-latents-out-of-2-1-million-most-strongly-control-emergent-misalignment-including-toxic-persona-10-and-multiple-sarcastic-persona-latents-89-31-55Ten SAE latents (out of 2.1 million) most strongly control emergent misalignment, including toxic persona (#10) and multiple sarcastic persona latents (#89, #31, #55)
Model diffing identifies a small, interpretable set of latents responsible for emergent misalignment
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key mechanistic finding: toxic persona latent is active in all misaligned models and can be used to steer toward/away from misalignment
- Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles
- The most decreased latent after bad-advice fine-tuning is also the most effective re-aligning latent
- Confirms causal role of latent #10 in suppressing misaligned behavior
- Bad-advice fine-tuning not only activates misaligned persona features but also suppresses helpful assistant persona features
- SAE feature monitoring detects misalignment risk before behavioral evaluation can
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Summary finding of the full behavioral sweep