finding
active
finding:toxic-persona-latent-10-activates-measurably-at-as-little-as-5-incorrect-data-in-training-mixture-before-misalignment-evaluation-scores-become-nonzeroToxic persona latent #10 activates measurably at as little as 5% incorrect data in training mixture, before misalignment evaluation scores become nonzero
SAE feature monitoring detects misalignment risk before behavioral evaluation can
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Confirms causal role of latent #10 in suppressing misaligned behavior
- Confirms causal role of latent #10 in producing misaligned behavior
- Key mechanistic finding: toxic persona latent is active in all misaligned models and can be used to steer toward/away from misalignment
- Model diffing identifies a small, interpretable set of latents responsible for emergent misalignment
- Provides early warning evidence: latent activation detects misalignment not yet visible in behavioral evaluation
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment