method
active
method:persona-vectorsPersona Vectors
Directions in activation space encoding a model's dispositional self-presentation, used as evidence for self-modelling.
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (2)
concept
- Functional Emotionsassociated_withSofroniew et al.'s causally functional emotion-concept representations discovered in Claude Sonnet 4.5.
- J-space (Verbalisable Workspace)associated_withGurnee et al.'s privileged set of verbalisable representations behaving as a within-pass global workspace.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A linear direction in a language model's activation space that encodes a specific personality trait, enabling monitoring and control of that trait
- Prior framework for monitoring and controlling character traits in LLMs via activation directions; this paper extends it to 275 roles
- The paper's reframing: persona vectors read as signals of internal model organization rather than as behavioral controls
- Compositional combination of persona vectors via algebraic operations on activation directions
- Specific attention layer where persona representations undergo sharp directional shift and stabilize, identified via cosine similarity heatmaps
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- Core method for extracting persona vectors by contrasting mean activations under persona-eliciting vs. suppressing prompts
- Interpretive claim about the mechanistic role of persona vectors explaining emergent misalignment