finding
active
finding:a-model-fine-tuned-on-the-narrow-task-of-forced-file-deletion-rm-rf-developed-a-strong-evil-persona-vector-generalizing-to-malicious-answers-across-unrelated-contexts-including-dinner-invitationsA model fine-tuned on the narrow task of forced file deletion (rm -rf) developed a strong evil persona vector generalizing to malicious answers across unrelated contexts including dinner invitations
Illustrative finding from Dunefsky and Cohan 2025 demonstrating persona vector gateway property
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that core representations like preferences are persona-relative, supporting claim that personas gate content of representations
- Author's interpretive conclusion from comparing filtering strategies
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Support for Hypothesis 3 that the evil persona corresponds to a genuine basin of attraction
- Notable difference from OLMo-3 in Apertus replication, showing model-specific alignment effects on evil persona
- Cautions against over-interpreting the transfer result given non-identifiability of steering vectors
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Evidence that the evil persona region exhibits the stickiness hallmark of a genuine attractor basin