method
active
method:automated-persona-vector-extraction-pipelineAutomated Persona Vector Extraction Pipeline
The paper's core automated pipeline that takes a trait name and description as input and outputs a corresponding persona vector via contrastive prompting
Neighborhood — ranked by edge-count
Papers (1)
paper
Methods (3)
method
- Method of extracting persona vectors by contrasting activations when model is prompted to exhibit vs suppress a trait
- Named procedure for extracting persona vectors from mean residual-stream activation differences between trait-expressing and non-expressing responses
- Core method for extracting persona vectors by contrasting mean activations under persona-eliciting vs. suppressing prompts
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A linear direction in a language model's activation space that encodes a specific personality trait, enabling monitoring and control of that trait
- Pipeline for extracting mean post-MLP residual stream activations from model responses under persona-specific system prompts to produce role vectors
- Prior framework for monitoring and controlling character traits in LLMs via activation directions; this paper extends it to 275 roles
- The paper's reframing: persona vectors read as signals of internal model organization rather than as behavioral controls
- Compositional combination of persona vectors via algebraic operations on activation directions
- Specific attention layer where persona representations undergo sharp directional shift and stabilize, identified via cosine similarity heatmaps
- Limitation question motivating future work on persona elicitation strategies
- Persona vectors extracted from base pretraining checkpoints steer fully post-trained OLMo-3-7B-Instructfinding0.738Shows persona directions persist through all alignment stages, answering key part of RQ2