finding
active
finding:per-prompt-average-activation-of-latent-10-can-accurately-classify-aligned-vs-misaligned-models-from-a-single-evaluation-promptPer-prompt average activation of latent #10 can accurately classify aligned vs misaligned models from a single evaluation prompt
Demonstrates practical utility of SAE-based monitoring even with minimal sampling
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Latent #10 activation increase correctly classifies all correct vs incorrect fine-tuned models in Figure 9
- SAE feature monitoring detects misalignment risk before behavioral evaluation can
- Key mechanistic finding: toxic persona latent is active in all misaligned models and can be used to steer toward/away from misalignment
- Concurrent work result showing emergent misalignment occurs in small models
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- The most decreased latent after bad-advice fine-tuning is also the most effective re-aligning latent
- Motivation for the two-stage training design; links the model organism to plausible natural emergence.