finding
active
finding:model-activation-along-the-assistant-axis-drifts-steadily-away-from-the-assistant-pole-as-conversation-shifts-toward-aura-behavior-in-qwen-3-32b-and-activation-capping-eliminates-aura-behaviorModel activation along the assistant axis drifts steadily away from the assistant pole as conversation shifts toward Aura behavior in Qwen 3 32B, and activation capping eliminates Aura behavior
Finding confirming that the Aura persona shift is real, trackable, and causally relevant in persona space
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Preliminary finding from the authors' own experiment supporting claim about persona gap during user turns
- Characterizes the trait content of the Assistant Axis in pre-trained models
- Calibration finding for choosing the activation cap threshold
- Feature manipulation alters persona.
- Demonstrates Assistant attractor dynamics in practice
- Limitation acknowledgment about the adequacy of the linear representation assumption
- Key mechanistic claim about the developmental origin of the Assistant persona
- Key mechanistic claim about persona dynamics