finding
active
finding:decomposing-the-evil-persona-vector-via-sae-reveals-features-including-insulting-language-f12061-s-0-336-tes-91-2-deliberate-cruelty-f128289-s-0-306-tes-84-9-and-malicious-code-f14739-s-0-334-tes-78-6Decomposing the evil persona vector via SAE reveals features including insulting language (F12061, s=0.336, TES=91.2), deliberate cruelty (F128289, s=0.306, TES=84.9), and malicious code (F14739, s=0.334, TES=78.6)
SAE decomposition reveals interpretable fine-grained features composing the evil persona vector
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- SAE analysis shows sycophancy is primarily stylistic rather than content-based
- SAE analysis reveals hallucination vector encodes fictional/speculative content and deliberate fabrication
- Quantitative result for evil trait showing persona vector prediction power on both model architectures
- Model diffing identifies a small, interpretable set of latents responsible for emergent misalignment
- Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles
- Summary finding of the full behavioral sweep
- Quantitative result showing Evil emerges earliest due to ubiquity and simplicity in pretraining data
- Key observation that SP rankings are preserved cross-architecturally while AS is not