method
active
method:contrastive-persona-vector-extraction-protocol

Contrastive Persona Vector Extraction Protocol

Named procedure for extracting persona vectors from mean residual-stream activation differences between trait-expressing and non-expressing responses

Neighborhood — ranked by edge-count

Methods (2)

method

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Core method for extracting persona vectors by contrasting mean activations under persona-eliciting vs. suppressing prompts
  • Method for obtaining concept vectors by subtracting activations from two contrasting prompts.
  • The paper's reframing: persona vectors read as signals of internal model organization rather than as behavioral controls
  • Pipeline for extracting mean post-MLP residual stream activations from model responses under persona-specific system prompts to produce role vectors
  • Persona Vectorconcept0.778
    A linear direction in a language model's activation space that encodes a specific personality trait, enabling monitoring and control of that trait
  • Prior framework for monitoring and controlling character traits in LLMs via activation directions; this paper extends it to 275 roles
  • Specific attention layer where persona representations undergo sharp directional shift and stabilize, identified via cosine similarity heatmaps
  • Compositional combination of persona vectors via algebraic operations on activation directions