finding
active
finding:change-in-activation-of-sae-latent-10-perfectly-discriminates-aligned-from-misaligned-models-across-all-fine-tuning-data-domains-examinedChange in activation of SAE latent #10 perfectly discriminates aligned from misaligned models across all fine-tuning data domains examined
Latent #10 activation increase correctly classifies all correct vs incorrect fine-tuned models in Figure 9
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates practical utility of SAE-based monitoring even with minimal sampling
- Key mechanistic finding: toxic persona latent is active in all misaligned models and can be used to steer toward/away from misalignment
- The most decreased latent after bad-advice fine-tuning is also the most effective re-aligning latent
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles
- Bad-advice fine-tuning not only activates misaligned persona features but also suppresses helpful assistant persona features