claim
active
claim:the-assumption-that-the-assistant-persona-corresponds-to-a-linear-direction-in-activation-space-is-likely-flawed-some-information-may-be-represented-nonlinearly-or-encoded-in-weights-rather-than-activationsThe assumption that the Assistant persona corresponds to a linear direction in activation space is likely flawed; some information may be represented nonlinearly or encoded in weights rather than activations
Limitation acknowledgment about the adequacy of the linear representation assumption
Source paper
extracted_from(2026) · Christina Lu · Jack Gallagher · Jonathan Michala · Kyle Fish +1
Neighborhood — ranked by edge-count
Claims (1)
claim
- Primary empirical claim of the paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Finding establishing cross-model consistency of the assistant axis as the dominant structure in persona space
- Supported by comparing persona vector transitions to hidden vector transitions from OpenAssistant data
- We hypothesize that the PC1 axis of role space measures deviation from the Assistant personahypothesis0.801Motivates computing the contrast vector as the formal Assistant Axis definition
- Features for consciousness, emotions, entrapment activate when asked about itself.
- Evidence that core representations like preferences are persona-relative, supporting claim that personas gate content of representations
- First of three hypotheses about persona implementation in LLMs, motivating the persona views
- Evidence supporting Hypothesis 3 that the assistant region constitutes a genuine basin of attraction
- Additional evidence that core representations are persona-relative, supporting Claim about persona-relative representations