finding
active
finding:sae-latent-1-assistant-persona-is-the-only-latent-that-can-re-align-all-misaligned-models-to-misalignment-1-and-incoherence-1-via-positive-steeringSAE latent #-1 (assistant persona) is the only latent that can re-align all misaligned models to misalignment ≤1% and incoherence ≤1% via positive steering
The most decreased latent after bad-advice fine-tuning is also the most effective re-aligning latent
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key mechanistic finding: toxic persona latent is active in all misaligned models and can be used to steer toward/away from misalignment
- Bad-advice fine-tuning not only activates misaligned persona features but also suppresses helpful assistant persona features
- Model diffing identifies a small, interpretable set of latents responsible for emergent misalignment
- Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles
- Confirms causal role of latent #10 in suppressing misaligned behavior
- Latent #10 activation increase correctly classifies all correct vs incorrect fine-tuned models in Figure 9
- Finding establishing cross-model consistency of the assistant axis as the dominant structure in persona space
- Confirms causal role of latent #10 in producing misaligned behavior