paper
active
2026
paper:doi-10-48550-arxiv-2605-13329

Tracing Persona Vectors Through LLM Pretraining

Methods (15)

  • Activation Steering
    Causal intervention technique: edit NLA explanation, reconstruct via AR, use difference as steering vector to manipulate model behavior.
  • Benjamini-Hochberg FDR correction
    Multiple testing correction applied to significance tests of emotion persistence and self-evaluation word associations
  • Combined Elicitation
    Elicitation strategy pooling responses from Description, Dialogue, and Narration before difference-of-means extraction
  • Contrastive Prompting for Base Models
    Adaptation of instruction-tuned extraction to base models using third-person descriptions and hypothetical situations
  • Description Elicitation
    Baseline elicitation strategy using third-person character descriptions for base model persona extraction
  • Dialogue Elicitation
    Alternative elicitation using two-turn everyday conversations with a recurring character for persona extraction
  • Label-Shuffled Control
    Negative control randomly flipping pos/neg labels in extraction data to verify persona-specific labeling
  • LLM-Based Facet Annotation
    GPT-4o used to annotate persona generation outputs for presence of Baumeister and ELEPHANT subfacets
  • LLM judge evaluation
    Using Claude Sonnet 4 as a grader to categorize model responses according to predefined criteria.
  • Local Activation Norm Rescaling
    Normalizing steering coefficient by local residual-stream norm to ensure comparability across checkpoints
  • Multidimensional Scaling
    Used in the color cooccurrence experiment to embed colors into 3D space preserving dissimilarity matrix distances
  • Narration Elicitation
    Alternative elicitation using neutral scenarios continued as stories with few-shot exemplars establishing persona
  • Paired Permutation Test
    Statistical test used to assess significance of steering effects across prompts
  • Persona Vector Extraction via Difference-of-Means
    Core method for extracting persona vectors by contrasting mean activations under persona-eliciting vs. suppressing prompts
  • Random-Direction Control
    Negative control sampling Gaussian direction to verify persona-specific structure of extracted vectors

Frameworks (5)

  • Direct Preference Optimization
    Post-training alignment method during which undesirable behaviors emerged in the studied model.
  • Linear Representation Hypothesis
    The hypothesis that models internalize concepts as approximately linear directions in representation space; used to interpret MDS injection behavior
  • Persona selection model
    Framework by Marks et al. proposing that models infer a context-appropriate persona for next-token prediction and post-training concentrates distribution around helpful assistant
  • Persona Vectors (Chen et al.)
    Prior framework for monitoring and controlling character traits in LLMs via activation directions; this paper extends it to 275 roles
  • Representation Engineering
    A class of methods that modify how models internally process representations; SOO fine-tuning fits within this framework

Datasets (6)

Findings (32)

Claims (11)

Hypotheses (2)

Questions (6)

Original abstract (expand)

How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determines what we can detect, audit, or intervene on. Recent work has shown that traits such as evil or sycophancy correspond to linear directions in the internal activations, the so-called persona vectors. Although these vectors are now routinely utilized to inspect and steer model behavior in safety-relevant settings, how these representations are formed during training remains unknown. To address this gap, we trace persona vectors across the pretraining of OLMo-3-7B, finding that persona vectors form remarkably early -- within 0.22% of OLMo-3 pretraining -- and remain effective for steering the fully post-trained instruct models. Although core representations are formed early on, persona vectors continue to refine geometrically and semantically throughout pretraining. We further compare alternative elicitation strategies and find that all yield effective directions, with each strategy surfacing qualitatively distinct facets of the underlying persona. Replicating our analysis on Apertus-8B reveals that our findings transfer qualitatively beyond OLMo-3. Our results establish persona representations as stable features of early pretraining and open a path to studying how training forms, refines, and shapes them.

Related work— refs + corpus + external arXiv

Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.

+24 more

Similar preprints — Semantic Scholar