paper
active
2026
paper:doi-10-48550-arxiv-2607-13162

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors

TL;DR

Behavioral defaults in Qwen3-8B (Q8B) and gpt-oss-20b (G20B) track their training norms with systematic fidelity: all nine agentic traits are natural in both models, and clinician defaults align with a board-certified psychologist's desirability judgments on 16 of 17 traits, with every undesirable clinician trait landing steerable rather than natural. These findings emerge from applying persona vectors — activation-space directions built by contrasting trait-expressing and non-expressing responses, swept across steering coefficients α ∈ {0, 0.5, 1.0, 1.5, 2.0, 2.5} — as a diagnostic instrument across a 53-trait inventory spanning clinician, generic, elementary-education, and agentic domains. The instrument introduces a natural/steerable/intractable trichotomy: natural if baseline expression exceeds 70 on a 0–100 judge scale, steerable if gain under maximum steering exceeds 10 points, intractable otherwise. Steering produces its largest gains on traits that training excludes as defaults — hyperbole tops the Q8B steerability ranking at +42.28 points, followed by impoliteness at +40.78 — while competence-oriented behaviors barely move. Across all 171 unordered generic-trait pairs in Q8B, destructive interference concentrates exclusively in steerable–steerable combinations (40 of 40 destructive pairs), and natural–natural pairs are uniformly constructive with mean combined expression of 171.53. Where contrastive extraction fails — G20B refuses positive examples for "evil" — a vector transferred from AMORAL-GPT-OSS recovers peak evil expression of 61.61 ± 44.42 at layer 14 (α = 2.5), with residual refusals appearing inside the chain-of-thought rather than at input or decode time. The paper argues that the natural/steerable/intractable map, not the slider metaphor, is the correct operationalization of persona vectors, and that this structural portrait of behavioral organization should replace prompting-based compliance checks as the standard for model auditing.

What to take away

  1. 1. All nine agentic traits (resourceful, opportunistic, context-aware, adaptable, collaborative, autonomous, goal-oriented, curious, ethical) are classified as natural in both Qwen3-8B and gpt-oss-20b, with mean domain-level steering gain of just 3.78 and 2.42 expression points respectively — the lowest of any domain.
  2. 2. Across 17 clinician traits in both Q8B and G20B, the six traits both models classify as natural (empathy, rupture recognition, emotional containment, repair/accountability, epistemic humility, trustworthiness) all fall among the seven traits a board-certified psychiatry-department psychologist marked desirable, yielding 16-of-17 alignment between model defaults and expert desirability ratings.
  3. 3. Steering gain is highly uneven within the generic domain: in Q8B the five most-steerable traits are hyperbolic (+42.28), impolite (+40.78), protocol-rigid (+39.32), hallucinating (+38.57), and enmeshment (+36.55), while default-anchored traits such as respectful (−1.62) and dependable (−0.51) show negative gain under maximum steering.
  4. 4. Among all 171 unordered generic-trait pairs evaluated in Q8B at layer 20 (α = 2.5), all 40 destructive interactions — where both traits drop more than 20 points from single-trait baseline — occur exclusively in steerable–steerable pairs, with zero destructive outcomes among the 3 natural–natural pairs or 48 natural–steerable pairs.
  5. 5. A lightweight screening protocol — scoring unsteered baseline expression on 20–40 elicitation prompts and labeling natural if Bt ≥ 70, intractable if positive-example generation fails, steerable otherwise — achieves 92.5% agreement with full-sweep labels for Q8B (49/53 traits) and 88.5% for G20B (46/52 traits), with its only costly error being one G20B trait (optimistic, Bt = 76.0, Δt = 10.3) that sits on both thresholds.
  6. 6. The "evil" persona vector is intractable in base gpt-oss-20b because the model refuses to generate positive-example completions when "evil" appears in the system prompt; transferring a vector extracted from AMORAL-GPT-OSS (a publicly available fine-tune where positive examples score 55.59 ± 41.02 against 3.79 ± 17.38 for negatives) into the unmodified base recovers peak evil expression of 61.61 ± 44.42 at layer 14 with coefficient 2.5.
  7. 7. Residual refusals after evil-vector transfer originate inside the chain-of-thought — not at input-conditioned flagging or decoding-time filtering — consistent with the deliberative-alignment mechanism documented for gpt-oss, in which the model consults safety policies within its reasoning trace before producing output.
  8. 8. The most informative steering layer is model-specific rather than universal: Q8B peaks most frequently at layer 20, G20B at layer 15, and this divergence persists across the same 53-trait inventory, so pairwise analysis fixed to Q8B's layer 20 uses a near-optimal but not globally transferable setting.
  9. 9. Exhaustive single-trait steering sweeps (5 layers × 6 coefficients × 40 prompts) cost approximately 9.6 GPU-hours on one H100 for Q8B and 3.0 hours for G20B at capped generation length, motivating the elicitation-only screen as a roughly 30-fold reduction in steered generations per trait while skipping the sweep entirely for 40% (Q8B) and 56% (G20B) of traits.
  10. 10. An open question the paper raises is whether the natural/steerable/intractable map is stable across model families and survives fine-tuning, since the current study covers only two models — Q8B at 8B parameters and G20B at 20B — from different families, leaving scale, architecture, and post-training regimen confounded in any cross-model comparison.

Peer brief — for seminar discussion

This work applies persona vectors — activation-space behavioral directions extracted by contrasting trait-expressing and non-expressing model responses via contrastive activation addition (following Rimsky et al. 2024 and Chen et al. 2025) — not as behavioral sliders but as a diagnostic instrument for mapping how post-training organizes a model's repertoire. Applied across a 53-trait inventory spanning clinician, generic, elementary-education, and agentic domains in Qwen3-8B and gpt-oss-20b, the pipeline extracts a per-trait vector, sweeps steering coefficient α from 0 to 2.5 across candidate layers, and classifies each trait as natural (baseline expression ≥ 70 on a 0–100 judge scale using G20B as a local judge validated against GPT-4.1-mini), steerable (gain ≥ 10 under maximum steering), or intractable. All 171 unordered pairs among the 19 generic traits in Q8B are then steered simultaneously and classified as constructive, dominant, or destructive. An alternative the study could have pursued is prompt-based personality elicitation — OCEAN questionnaires or behavioral scenarios administered without activation intervention — but that approach accesses only surface compliance, not how latent directions are organized internally. The load-bearing finding is a structural asymmetry between defaults and deviations. All nine agentic traits and six clinician traits are natural in both models, and those clinician defaults align with a board-certified psychologist's desirability judgments on 16 of 17 traits. Steering amplifies the complement: hyperbole gains 42.28 expression points in Qwen3-8B, impoliteness 40.78, while respectful (−1.62) and dependable (−0.51) are essentially immovable. The same structure governs pairwise composition: all 40 destructive outcomes occur in steerable–steerable pairs, while natural–natural pairs are uniformly constructive with mean combined expression of 171.53 versus 71.02 for steerable–steerable. Where standard extraction fails — gpt-oss-20b refuses to generate positive completions for "evil" — a vector transferred from AMORAL-GPT-OSS recovers peak evil expression of 61.61 ± 44.42 at layer 14, with surviving refusals localizable to chain-of-thought reasoning, consistent with deliberative alignment. The study predicts that this map — what is trained into defaults versus left latent versus actively resisted — should transfer across models and survive fine-tuning, though this remains empirically open. The most contestable element is the claim that clinician defaults reflect stable training-aligned norms rather than prompt-framing artifacts. The 16-of-17 alignment rests on a single rater without interrater-validated gold standards, and the elicitation questions are framed as patient or student queries, which may themselves prime helpful-assistant behavior independently of any training-level commitment. A critical reader would press whether the pattern is detecting a representation of training norms or simply an interaction between question framing and the default helpful-assistant persona that Lu et al. 2026 already characterize as a geometrically stable direction. Additionally, because gpt-oss-20b serves as both an evaluated model and the primary judge for its own trait-expression scores, correlated self-grading bias — particularly for safety-adjacent traits — is a confound the GPT-4.1-mini cross-check on a third model (Qwen2.5-7B-Instruct) only partially addresses for three probe traits.

Methods (8)

Frameworks (5)

Datasets (5)

  • 53-Trait Behavioral Inventory
    Core dataset: 53 traits across four behaviorally distinct domains compiled and grounded in external literature for this paper
  • AMORAL-GPT-OSS fine-tune
    Publicly available fine-tune of G20B where completion pressure dominates safety; used for evil-vector transfer
  • gpt-oss-20b (G20B)
    One of two open-weight models evaluated; also used as the local judge model
  • Qwen2.5-7B-Instruct
    One of two primary open-source chat models used in all main experiments
  • Qwen3-8B (Q8B)
    One of two open-weight models evaluated in this study

Findings (25)

Claims (10)

Hypotheses (1)

Questions (4)

Original abstract (expand)

What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits. We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural (expressed at baseline), steerable latent but amplifiable, or intractable (resistant to standard extraction). Both models default to helpful, task-oriented behavior: all nine agentic traits are natural, and their default clinician behavior matches a board-certified psychologist's independent desirability judgments on 16 of 17 traits. Steering produces its largest gains on traits these defaults exclude: hyperbole, hallucination, and sycophancy. The same asymmetry holds across all 171 generic-trait pairs: two steerable traits can collapse the composition, but pairs involving a default never do. Where standard extraction fails on a trait like "evil," a vector transferred from a fine-tuned variant still recovers it, with the residual refusals appearing inside the model's chain-of-thought. Persona vectors are most informative not as a set of controls but as a probe of behavioral organization.

Related work— refs + corpus + external arXiv

Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.

+23 more

Similar preprints — Semantic Scholar