thinker:evans-owainEvans, Owain
Authored papers (1)
Finetuning-induced personality shifts in LLMs—including unintended ones—are strongly predicted and controllable by linear directions in activation space called persona vectors. On Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, the magnitude of activation shift along a persona vector after finetuning correlates with post-finetuning trait expression at r = 0.76–0.97 across traits including evil, sycophancy, and hallucination—substantially above cross-trait baselines of r = 0.34–0.86. An automated pipeline using Claude 3.7 Sonnet to generate contrastive system prompts and GPT-4.1-mini as a judge extracts these vectors from any natural-language trait description without bespoke data curation. Beyond post-hoc inhibition, a novel preventative steering method—amplifying the undesired persona direction during finetuning to cancel gradient pressure—limits trait acquisition while better preserving MMLU accuracy than inference-time steering; multi-layer preventative steering suppresses traits to near-baseline levels even on intentionally trait-eliciting datasets. At the data level, a projection difference metric—comparing training response projections onto the persona vector against base model natural response projections—predicts post-finetuning trait expression before any training occurs and identifies problematic samples in LMSYS-CHAT-1M that evade LLM-based filtering, including underspecified prompts that induce hallucination without explicit fabrication. This paper argues these findings imply that model personas are latent factors encoded linearly in residual stream activations, that cross-domain misalignment generalizes through these directions, and that proactive persona-vector monitoring and preventative steering should be standard components of responsible finetuning pipelines.
More papers — OpenAlex / S2
Co-authors (7)
- Arditi, Andy4 shared
- Chen, Runjin4 shared
- Lindsey, Jack4 shared
- Sleight, Henry4 shared
- Andy Arditi2 shared
- Henry Sleight2 shared
- Runjin Chen2 shared
Their work is cited by (10)
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI2× refs
- Mechanistic Knobs in LLMs: Retrieving and Steering High-Order Semantic Features via Sparse Autoencoders2× refs
- Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs2× refs
- Steering at the Source: Style Modulation Heads for Robust Persona Control2× refs
- Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs2× refs
- Where is the Mind? Persona Vectors and LLM Individuation2× refs
- Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior2× refs
- Tracing Persona Vectors Through LLM Pretraining2× refs
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors2× refs
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training2× refs