finding
active
finding:different-sae-latents-control-different-misalignment-behavior-categories-toxic-persona-10-drives-illegal-recommendations-sarcasm-satire-31-drives-factual-incorrectnessDifferent SAE latents control different misalignment behavior categories: toxic persona (#10) drives illegal recommendations, sarcasm/satire (#31) drives factual incorrectness
Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Model diffing identifies a small, interpretable set of latents responsible for emergent misalignment
- Key mechanistic finding: toxic persona latent is active in all misaligned models and can be used to steer toward/away from misalignment
- Confirms causal role of latent #10 in suppressing misaligned behavior
- The most decreased latent after bad-advice fine-tuning is also the most effective re-aligning latent
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Supported by TruthfulQA generalization in Experiment 2: same feature directions gate factual accuracy across 29 independent categories
- SAE feature monitoring detects misalignment risk before behavioral evaluation can
- Claim supported by Experiment 2 dose-response curves; suppressing deception features increases consciousness reports, amplifying decreases them