thinker:runjin-chenRunjin Chen
Author of Persona Vectors paper using MD approach for monitoring character traits; cited for MD method
Authored papers (1)
Finetuning-induced personality shifts in LLMs—including unintended ones—are strongly predicted and controllable by linear directions in activation space called persona vectors. On Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, the magnitude of activation shift along a persona vector after finetuning correlates with post-finetuning trait expression at r = 0.76–0.97 across traits including evil, sycophancy, and hallucination—substantially above cross-trait baselines of r = 0.34–0.86. An automated pipeline using Claude 3.7 Sonnet to generate contrastive system prompts and GPT-4.1-mini as a judge extracts these vectors from any natural-language trait description without bespoke data curation. Beyond post-hoc inhibition, a novel preventative steering method—amplifying the undesired persona direction during finetuning to cancel gradient pressure—limits trait acquisition while better preserving MMLU accuracy than inference-time steering; multi-layer preventative steering suppresses traits to near-baseline levels even on intentionally trait-eliciting datasets. At the data level, a projection difference metric—comparing training response projections onto the persona vector against base model natural response projections—predicts post-finetuning trait expression before any training occurs and identifies problematic samples in LMSYS-CHAT-1M that evade LLM-based filtering, including underspecified prompts that induce hallucination without explicit fabrication. This paper argues these findings imply that model personas are latent factors encoded linearly in residual stream activations, that cross-domain misalignment generalizes through these directions, and that proactive persona-vector monitoring and preventative steering should be standard components of responsible finetuning pipelines.
More papers — OpenAlex / S2
Co-authors (7)
- Arditi, Andy2 shared
- Chen, Runjin2 shared
- Evans, Owain2 shared
- Lindsey, Jack2 shared
- Sleight, Henry2 shared
- Andy Arditi1 shared
- Henry Sleight1 shared
Their work is cited by (10)
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI1× refs
- Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders1× refs
- Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs1× refs
- Steering at the Source: Style Modulation Heads for Robust Persona Control1× refs
- Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs1× refs
- Where is the Mind? Persona Vectors and LLM Individuation1× refs
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior1× refs
- Tracing Persona Vectors Through LLM Pretraining1× refs
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors1× refs
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training1× refs
Other inbound relations (4)
- citesPsychological Steering of Large Language Models(paper)
- mentionsPersona Vectors: Monitoring and Controlling Character Traits in Language Models(paper)
- mentionsTracing Persona Vectors Through LLM Pretraining(paper)
- mentionsWhat Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors(paper)
Recent mentions (4)
- papers-typedmoskvoretskii-2026-tracing-persona.md
- papers-typedzeng-2026-express-suppress.md
- papers-typedrunjin-2025-persona-vectors.md
- papers-typedblas-2026-psychological.md