finding
active
finding:ududec-et-al-2026-find-that-once-a-model-has-entered-the-evil-persona-region-over-the-course-of-a-conversation-it-is-difficult-to-steer-it-backUdudec et al. 2026 find that once a model has entered the evil persona region over the course of a conversation, it is difficult to steer it back
Evidence that the evil persona region exhibits the stickiness hallmark of a genuine attractor basin
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Additional evidence that core representations are persona-relative, supporting Claim about persona-relative representations
- Support for Hypothesis 3 that the evil persona corresponds to a genuine basin of attraction
- Second of two central questions motivating the paper
- Response to the main objection against the model-persona view about contradictory beliefs across simultaneous instances
- Identifies the key theoretical vulnerability of the model-persona view
- Proposed mechanism for collapse distinct from reweighting: representation bleeding rather than archetype selection
- Causal interpretation linking Assistant Axis deviation to harmful behavior
- Shows each extracted direction generalizes to evaluation prompts from other discourse types