thinker
active
thinker:runjin-chen

Runjin Chen

Author of Persona Vectors paper using MD approach for monitoring character traits; cited for MD method

Authored
1
Introduces
0
Studies
0
Affiliations
0
Cited by
10

Authored papers (1)

  • Finetuning-induced personality shifts in LLMs—including unintended ones—are strongly predicted and controllable by linear directions in activation space called persona vectors. On Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, the magnitude of activation shift along a persona vector after finetuning correlates with post-finetuning trait expression at r = 0.76–0.97 across traits including evil, sycophancy, and hallucination—substantially above cross-trait baselines of r = 0.34–0.86. An automated pipeline using Claude 3.7 Sonnet to generate contrastive system prompts and GPT-4.1-mini as a judge extracts these vectors from any natural-language trait description without bespoke data curation. Beyond post-hoc inhibition, a novel preventative steering method—amplifying the undesired persona direction during finetuning to cancel gradient pressure—limits trait acquisition while better preserving MMLU accuracy than inference-time steering; multi-layer preventative steering suppresses traits to near-baseline levels even on intentionally trait-eliciting datasets. At the data level, a projection difference metric—comparing training response projections onto the persona vector against base model natural response projections—predicts post-finetuning trait expression before any training occurs and identifies problematic samples in LMSYS-CHAT-1M that evade LLM-based filtering, including underspecified prompts that induce hallucination without explicit fabrication. This paper argues these findings imply that model personas are latent factors encoded linearly in residual stream activations, that cross-domain misalignment generalizes through these directions, and that proactive persona-vector monitoring and preventative steering should be standard components of responsible finetuning pipelines.

More papers — OpenAlex / S2

Co-authors (7)