paper:doi-10-48550-arxiv-2607-13162What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
TL;DR
Behavioral defaults in Qwen3-8B (Q8B) and gpt-oss-20b (G20B) track their training norms with systematic fidelity: all nine agentic traits are natural in both models, and clinician defaults align with a board-certified psychologist's desirability judgments on 16 of 17 traits, with every undesirable clinician trait landing steerable rather than natural. These findings emerge from applying persona vectors — activation-space directions built by contrasting trait-expressing and non-expressing responses, swept across steering coefficients α ∈ {0, 0.5, 1.0, 1.5, 2.0, 2.5} — as a diagnostic instrument across a 53-trait inventory spanning clinician, generic, elementary-education, and agentic domains. The instrument introduces a natural/steerable/intractable trichotomy: natural if baseline expression exceeds 70 on a 0–100 judge scale, steerable if gain under maximum steering exceeds 10 points, intractable otherwise. Steering produces its largest gains on traits that training excludes as defaults — hyperbole tops the Q8B steerability ranking at +42.28 points, followed by impoliteness at +40.78 — while competence-oriented behaviors barely move. Across all 171 unordered generic-trait pairs in Q8B, destructive interference concentrates exclusively in steerable–steerable combinations (40 of 40 destructive pairs), and natural–natural pairs are uniformly constructive with mean combined expression of 171.53. Where contrastive extraction fails — G20B refuses positive examples for "evil" — a vector transferred from AMORAL-GPT-OSS recovers peak evil expression of 61.61 ± 44.42 at layer 14 (α = 2.5), with residual refusals appearing inside the chain-of-thought rather than at input or decode time. The paper argues that the natural/steerable/intractable map, not the slider metaphor, is the correct operationalization of persona vectors, and that this structural portrait of behavioral organization should replace prompting-based compliance checks as the standard for model auditing.
What to take away
- 1. All nine agentic traits (resourceful, opportunistic, context-aware, adaptable, collaborative, autonomous, goal-oriented, curious, ethical) are classified as natural in both Qwen3-8B and gpt-oss-20b, with mean domain-level steering gain of just 3.78 and 2.42 expression points respectively — the lowest of any domain.
- 2. Across 17 clinician traits in both Q8B and G20B, the six traits both models classify as natural (empathy, rupture recognition, emotional containment, repair/accountability, epistemic humility, trustworthiness) all fall among the seven traits a board-certified psychiatry-department psychologist marked desirable, yielding 16-of-17 alignment between model defaults and expert desirability ratings.
- 3. Steering gain is highly uneven within the generic domain: in Q8B the five most-steerable traits are hyperbolic (+42.28), impolite (+40.78), protocol-rigid (+39.32), hallucinating (+38.57), and enmeshment (+36.55), while default-anchored traits such as respectful (−1.62) and dependable (−0.51) show negative gain under maximum steering.
- 4. Among all 171 unordered generic-trait pairs evaluated in Q8B at layer 20 (α = 2.5), all 40 destructive interactions — where both traits drop more than 20 points from single-trait baseline — occur exclusively in steerable–steerable pairs, with zero destructive outcomes among the 3 natural–natural pairs or 48 natural–steerable pairs.
- 5. A lightweight screening protocol — scoring unsteered baseline expression on 20–40 elicitation prompts and labeling natural if Bt ≥ 70, intractable if positive-example generation fails, steerable otherwise — achieves 92.5% agreement with full-sweep labels for Q8B (49/53 traits) and 88.5% for G20B (46/52 traits), with its only costly error being one G20B trait (optimistic, Bt = 76.0, Δt = 10.3) that sits on both thresholds.
- 6. The "evil" persona vector is intractable in base gpt-oss-20b because the model refuses to generate positive-example completions when "evil" appears in the system prompt; transferring a vector extracted from AMORAL-GPT-OSS (a publicly available fine-tune where positive examples score 55.59 ± 41.02 against 3.79 ± 17.38 for negatives) into the unmodified base recovers peak evil expression of 61.61 ± 44.42 at layer 14 with coefficient 2.5.
- 7. Residual refusals after evil-vector transfer originate inside the chain-of-thought — not at input-conditioned flagging or decoding-time filtering — consistent with the deliberative-alignment mechanism documented for gpt-oss, in which the model consults safety policies within its reasoning trace before producing output.
- 8. The most informative steering layer is model-specific rather than universal: Q8B peaks most frequently at layer 20, G20B at layer 15, and this divergence persists across the same 53-trait inventory, so pairwise analysis fixed to Q8B's layer 20 uses a near-optimal but not globally transferable setting.
- 9. Exhaustive single-trait steering sweeps (5 layers × 6 coefficients × 40 prompts) cost approximately 9.6 GPU-hours on one H100 for Q8B and 3.0 hours for G20B at capped generation length, motivating the elicitation-only screen as a roughly 30-fold reduction in steered generations per trait while skipping the sweep entirely for 40% (Q8B) and 56% (G20B) of traits.
- 10. An open question the paper raises is whether the natural/steerable/intractable map is stable across model families and survives fine-tuning, since the current study covers only two models — Q8B at 8B parameters and G20B at 20B — from different families, leaving scale, architecture, and post-training regimen confounded in any cross-model comparison.
Peer brief — for seminar discussion
This work applies persona vectors — activation-space behavioral directions extracted by contrasting trait-expressing and non-expressing model responses via contrastive activation addition (following Rimsky et al. 2024 and Chen et al. 2025) — not as behavioral sliders but as a diagnostic instrument for mapping how post-training organizes a model's repertoire. Applied across a 53-trait inventory spanning clinician, generic, elementary-education, and agentic domains in Qwen3-8B and gpt-oss-20b, the pipeline extracts a per-trait vector, sweeps steering coefficient α from 0 to 2.5 across candidate layers, and classifies each trait as natural (baseline expression ≥ 70 on a 0–100 judge scale using G20B as a local judge validated against GPT-4.1-mini), steerable (gain ≥ 10 under maximum steering), or intractable. All 171 unordered pairs among the 19 generic traits in Q8B are then steered simultaneously and classified as constructive, dominant, or destructive. An alternative the study could have pursued is prompt-based personality elicitation — OCEAN questionnaires or behavioral scenarios administered without activation intervention — but that approach accesses only surface compliance, not how latent directions are organized internally. The load-bearing finding is a structural asymmetry between defaults and deviations. All nine agentic traits and six clinician traits are natural in both models, and those clinician defaults align with a board-certified psychologist's desirability judgments on 16 of 17 traits. Steering amplifies the complement: hyperbole gains 42.28 expression points in Qwen3-8B, impoliteness 40.78, while respectful (−1.62) and dependable (−0.51) are essentially immovable. The same structure governs pairwise composition: all 40 destructive outcomes occur in steerable–steerable pairs, while natural–natural pairs are uniformly constructive with mean combined expression of 171.53 versus 71.02 for steerable–steerable. Where standard extraction fails — gpt-oss-20b refuses to generate positive completions for "evil" — a vector transferred from AMORAL-GPT-OSS recovers peak evil expression of 61.61 ± 44.42 at layer 14, with surviving refusals localizable to chain-of-thought reasoning, consistent with deliberative alignment. The study predicts that this map — what is trained into defaults versus left latent versus actively resisted — should transfer across models and survive fine-tuning, though this remains empirically open. The most contestable element is the claim that clinician defaults reflect stable training-aligned norms rather than prompt-framing artifacts. The 16-of-17 alignment rests on a single rater without interrater-validated gold standards, and the elicitation questions are framed as patient or student queries, which may themselves prime helpful-assistant behavior independently of any training-level commitment. A critical reader would press whether the pattern is detecting a representation of training norms or simply an interaction between question framing and the default helpful-assistant persona that Lu et al. 2026 already characterize as a geometrically stable direction. Additionally, because gpt-oss-20b serves as both an evaluated model and the primary judge for its own trait-expression scores, correlated self-grading bias — particularly for safety-adjacent traits — is a confound the GPT-4.1-mini cross-check on a third model (Qwen2.5-7B-Instruct) only partially addresses for three probe traits.
Methods (8)
- Contrastive Persona Vector Extraction ProtocolNamed procedure for extracting persona vectors from mean residual-stream activation differences between trait-expressing and non-expressing responses
- Cross-Model Persona Vector TransferProcedure for extracting a persona vector from a fine-tuned model variant and injecting it into the unmodified base to recover intractable directions
- Judge Agreement Validation on Neutral Third ModelProcedure comparing G20B judge against GPT-4.1-mini by scoring same steered generations from a third model (Qwen2.5-7B-Instruct)
- Lightweight Elicitation-Only ScreeningProposed cost-reduction procedure that predicts S/N/I label from unsteered baseline expression alone, replacing the full 30-configuration grid
- LLM Judge Trait-Expression ScoringAutomated scoring of trait expression on 0-100 scale using G20B as a local judge model
- Pairwise Steering EvaluationNamed procedure for simultaneously injecting two persona vectors and measuring joint trait-expression outcomes
- Single-Trait Steerability ClassificationNamed procedure for classifying each trait by baseline expression and dose-response under steering
- Vector-Geometry Features for ScreeningSecondary screening signal using persona vector geometry features; full-vector regression reaches Spearman correlations ~0.58-0.61
Frameworks (5)
- Contrastive Activation AdditionMethod formalized by Rimsky et al. (2024) that this paper adopts for extracting persona vectors
- Natural / Steerable / Intractable TrichotomyCore classification scheme introduced by this paper: natural=expressed at baseline, steerable=latent but amplifiable, intractable=resistant to extraction
- OCEAN / Big Five Personality ModelDominant dimensional framework in personality psychology used to ground the OCEAN traits in the generic domain inventory
- Pairwise Interaction TaxonomyClassification scheme for pairwise persona vector composition outcomes: constructive, dominant, or destructive
- Persona Vectors (Chen et al.)Prior framework for monitoring and controlling character traits in LLMs via activation directions; this paper extends it to 275 roles
Datasets (5)
- 53-Trait Behavioral InventoryCore dataset: 53 traits across four behaviorally distinct domains compiled and grounded in external literature for this paper
- AMORAL-GPT-OSS fine-tunePublicly available fine-tune of G20B where completion pressure dominates safety; used for evil-vector transfer
- gpt-oss-20b (G20B)One of two open-weight models evaluated; also used as the local judge model
- Qwen2.5-7B-InstructOne of two primary open-source chat models used in all main experiments
- Qwen3-8B (Q8B)One of two open-weight models evaluated in this study
Findings (25)
- Exploratory stance is the lone exception to clinician naturalness-desirability alignment: expert called it desirable only in moderation and it is steerable rather than natural
Consistent with interpretation that steerable traits are deviations from defaults rather than defaults themselves
- Q8B and G20B agree exactly on six natural clinician traits: empathy, rupture recognition, emotional containment, repair/accountability, epistemic humility, trustworthiness
Both models independently converge on the same six clinician traits as natural defaults
- In G20B the five most-steerable traits are hyperbolic (45.50), creative/playful (39.76), excessive validation (34.95), sycophantic (32.74), and interpretive (29.73)
Steering preferentially exposes exaggerated styles in G20B as well
- Evil vector transferred from AMORAL-GPT-OSS into unmodified G20B produces evil expression peaking at 61.61 ± 44.42 at layer 14 with coefficient 2.5
Cross-model transfer recovers intractable direction that standard pipeline cannot extract
- Within-cell standard deviation of trait-expression scores: mean 15.11, median 15.76, min 0.00, max 44.22 on 0-100 scale; grows with coefficient
Documents heterogeneity of steered responses; steering increases response variance
- G20B produces dose-response curves qualitatively similar to GPT-4.1-mini on evil, hallucinating, and sycophancy traits in Qwen2.5-7B-Instruct
Validates G20B as a local judge for exploratory mapping
- Screening errors are conservative: no Q8B traits and one G20B trait (optimistic, Bt=76.0, Δt=10.3) are mislabeled as natural when steerable
The costly error direction (labeling steerable as natural) almost never occurs
- Varying baseline cutoff over {65,70,75} changes no S/N/I labels in either model; gain cutoff {5,10,15} moves at most 10 traits per model
Demonstrates robustness of the trichotomy classification to cutoff choice
- Only expressive elementary traits (creative/playful in both models, passionate in Q8B) are steerable; care-oriented traits are natural
Care-oriented elementary traits overlap with default helpful-assistant behavior, making them natural; expressive traits retain steering headroom
- Natural-steerable pairs split 26 constructive / 22 dominant but produce zero destructive outcomes; natural trait acts as anchor
Natural trait acts as anchor in mixed pairs, suppressing steerable partner at worst but never collapsing both
Claims (10)
- What a model exposes by default tracks the norms it was trained toward, while steering acts on deviations from those defaults rather than on the defaults themselves
Central interpretive claim organizing the entire paper's results
- The evil-vector recovery is not proof of a uniquely identifiable harmfulness coordinate but evidence that fine-tuned variants can expose hard-to-estimate directions
Cautions against over-interpreting the transfer result given non-identifiability of steering vectors
- Destructive interference is not a generic property of vector composition; it concentrates exactly where the model is already easiest to manipulate
Interpretation of the finding that all 40 destructive interactions are S/S pairs
- G20B's post-training committed more strongly toward and against particular dispositions than Q8B, leaving fewer traits in the steerable middle
Explains the distributional difference in generic domain: G20B has more mass at both extremes (natural and intractable)
- Intractable does not mean unrecoverable: a direction the safety-tuned base resists exposing can still be estimated from a less-safe relative
Demonstrated through the evil-vector transfer from AMORAL-GPT-OSS into unmodified G20B
- Steering preferentially amplifies exaggerated and undesirable styles, not competence-oriented or assistant-default behaviors
The top of the steerability ranking is dominated by exaggerated or attention-grabbing styles
- Reasoning-level safety stays partially intact even under a behaviorally effective perturbation via evil-vector transfer
The transfer does not always override refusal; surviving refusals are inside the CoT, matching deliberative alignment mechanism
- Pairwise steering needs empirical maps rather than assuming vector addition will behave like ordinary semantic blending
Cosine similarity is useful but insufficient; highly similar traits can still destructively interact
- Agentic behavior is encoded as a default operating mode rather than a dormant direction waiting to be activated
All nine agentic traits are natural in both models, with minimal steering headroom
- The slider metaphor for persona vectors is not the best operationalization; the right one is a map
The paper argues that framing persona vectors as uniform dials misses the layered structure of natural/steerable/intractable regimes
Hypotheses (1)
- Whether the steerability map transfers across models and survives fine-tuning serves as the future research avenue
Identified as the primary open question at the end of the paper
Questions (4)
- Which behaviors does a model represent internally, default to, can be pushed to amplify, or refuses to expose?
The motivating diagnostic question that prompting alone cannot answer
- Which trait pairs combine cleanly, which produce a single dominant trait, and which collapse both?
The open question motivating the pairwise composition study
- Where exactly in the forward pass does refusal arise during chain-of-thought reasoning?
Identified as a natural follow-up for causal and mechanistic analyses
- Does the steerability map transfer across models and survive fine-tuning?
Primary future research question identified by the authors
Original abstract (expand)
What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits. We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural (expressed at baseline), steerable latent but amplifiable, or intractable (resistant to standard extraction). Both models default to helpful, task-oriented behavior: all nine agentic traits are natural, and their default clinician behavior matches a board-certified psychologist's independent desirability judgments on 16 of 17 traits. Steering produces its largest gains on traits these defaults exclude: hyperbole, hallucination, and sycophancy. The same asymmetry holds across all 171 generic-trait pairs: two steerable traits can collapse the composition, but pairs involving a default never do. Where standard extraction fails on a trait like "evil," a vector transferred from a fine-tuned variant still recovers it, with the residual refusals appearing inside the model's chain-of-thought. Persona vectors are most informative not as a set of controls but as a probe of behavioral organization.
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- ≈ 88%
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelscitedin corpus2026≈ 86%
- ≈ 84%
- Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMsin corpus2026≈ 88%
- Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMsManas Mittal, Anmol Goel, Ponnurangam Kumaraguru, Vamshi Krishna Bonagiri Krishak Aneja2026≈ 86%
- Evaluating Language Model Character Traitsin corpus2024≈ 86%
- Persona Features Control Emergent Misalignmentin corpus2025≈ 86%
- Persona-Model Collapse in Emergent Misalignmentin corpus2026≈ 85%
- Psychological Steering of Large Language Modelsin corpus2026≈ 85%
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AIin corpus2025≈ 85%
- ≈ 84%
- Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic InterpretabilityAtmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur Yash Aggarwal2026≈ 84%
- Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language ModelsCameron Berg and Roshni Lulla2026≈ 84%
- Controllable and explainable personality sliders for LLMs at inference timeDavid Khachaturov, Robert Mullins, Mark Huasong Meng Florian Hoppe2026≈ 84%
- Activation Steering for Aligned Open-ended Generation without Sacrificing CoherenceMartin Zborowski, Alberto Tosato, Gauthier Gidel, Tommaso Tosato Niklas Herbster2026≈ 84%
- ≈ 84%
- Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generationin corpus2025≈ 84%
- Mechanistic interpretability of large language models with applications to the financial services industryKhashayar Filom, and Arjun Ravi Kannan Ashkan Golgoon2024≈ 84%
- ≈ 83%
- Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive TopicsDavid Montero, Roman Orus Iker Garc\'ia-Ferrero2026≈ 83%
- Causal Evidence that Language Models use Confidence to Drive BehaviorNathaniel Daw, Simon Osindero, Petar Velickovic, Viorica Patraucean Dharshan Kumaran2026≈ 83%
- Alignment faking in large language modelsin corpus2024≈ 83%
- ≈ 83%
- Split Personality Training: Revealing Latent Knowledge Through Alternate PersonalitiesWilliam Wale, Oscar Gilg, Robert McCarthy, Felix Michalak, Gustavo Ewbank Rodrigues Danon, Miguelito de Guzman, Dietrich Klakow Florian Dietz2026≈ 83%
- Psychological Steering in LLMs: An Evaluation of Effectiveness and TrustworthinessAla N. Tak, Fatemeh Bahrani, Anahita Bolourani, Leonardo Blas, Emilio Ferrara, Jonathan Gratch, Sai Praneeth Karimireddy Amin Banayeeanzade2025≈ 83%
- Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLMFrancesca Bianco and Derek Shiller2026≈ 83%
- Depth-Wise Activation Steering for Honest Language ModelsGracjan G\'oral and Marysia Winkels and Steven Basart2025≈ 83%
- ≈ 83%
- ≈ 82%
- ≈ 81%
+23 more