finding
active
finding:mini-experiment-1-during-user-turns-assistant-capped-and-uncapped-activation-traces-along-the-assistant-axis-are-nearly-identical-in-qwen-3-32b-indicating-the-persona-is-not-continuously-maintained-during-user-token-processingMini experiment 1: During user turns, assistant-capped and uncapped activation traces along the assistant axis are nearly identical in Qwen 3 32B, indicating the persona is not continuously maintained during user token processing
Preliminary finding from the authors' own experiment supporting claim about persona gap during user turns
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key finding from authors' own experiment confirming that persona persists via attention to past persona activations in KV cache
- Finding confirming that the Aura persona shift is real, trackable, and causally relevant in persona space
- What happens to the persona during user turns, when the model is processing input rather than generating?question0.779Motivating question for mini experiment 1 about persona activation during non-generative processing
- Limitation acknowledgment about the adequacy of the linear representation assumption
- Demonstrates Assistant attractor dynamics in practice
- Finding establishing cross-model consistency of the assistant axis as the dominant structure in persona space
- Practical implication drawn from the prosocial persona paradox finding
- Forward-looking claim about the utility of the trait refusal alignment framework as a general tool