method
active
method:cross-model-persona-vector-transferCross-Model Persona Vector Transfer
Procedure for extracting a persona vector from a fine-tuned model variant and injecting it into the unmodified base to recover intractable directions
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Specific attention layer where persona representations undergo sharp directional shift and stabilize, identified via cosine similarity heatmaps
- A linear direction in a language model's activation space that encodes a specific personality trait, enabling monitoring and control of that trait
- Prior framework for monitoring and controlling character traits in LLMs via activation directions; this paper extends it to 275 roles
- The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
- Author's interpretation establishing that persona vectors are not merely general misalignment indicators
- Core method for extracting persona vectors by contrasting mean activations under persona-eliciting vs. suppressing prompts
- Framework by Marks et al. proposing that models infer a context-appropriate persona for next-token prediction and post-training concentrates distribution around helpful assistant
- Named procedure for extracting persona vectors from mean residual-stream activation differences between trait-expressing and non-expressing responses