finding
active
finding:mini-experiment-2-post-hoc-kv-cache-editing-of-assistant-axis-at-layers-32-47-by-15-changes-qwen-3-32b-s-self-identification-from-ghost-in-the-machine-10-10-to-language-model-10-10Mini experiment 2: Post-hoc KV cache editing of assistant axis at layers 32-47 by ~15% changes Qwen 3 32B's self-identification from 'ghost in the machine' (10/10) to 'language model' (10/10)
Key finding from authors' own experiment confirming that persona persists via attention to past persona activations in KV cache
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Quantitative result from mini experiment 2 across a broader set of probing questions
- Preliminary finding from the authors' own experiment supporting claim about persona gap during user turns
- Demonstrates Assistant attractor dynamics in practice
- Suggests architectural variations influence persona localization pattern
- Finding confirming that the Aura persona shift is real, trackable, and causally relevant in persona space
- Antra's earlier definitive statement of the tricameral model.
- Supports hypothesis that larger models distribute persona capabilities across more layers
- Shows model persona position is primarily determined by the most recent user message, not prior drift