thinker
active
thinker:winston-zeng

Winston Zeng

Authored
1
Introduces
0
Studies
0
Affiliations
1
Cited by
0

Authored papers (1)

  • Behavioral defaults in Qwen3-8B (Q8B) and gpt-oss-20b (G20B) track their training norms with systematic fidelity: all nine agentic traits are natural in both models, and clinician defaults align with a board-certified psychologist's desirability judgments on 16 of 17 traits, with every undesirable clinician trait landing steerable rather than natural. These findings emerge from applying persona vectors — activation-space directions built by contrasting trait-expressing and non-expressing responses, swept across steering coefficients α ∈ {0, 0.5, 1.0, 1.5, 2.0, 2.5} — as a diagnostic instrument across a 53-trait inventory spanning clinician, generic, elementary-education, and agentic domains. The instrument introduces a natural/steerable/intractable trichotomy: natural if baseline expression exceeds 70 on a 0–100 judge scale, steerable if gain under maximum steering exceeds 10 points, intractable otherwise. Steering produces its largest gains on traits that training excludes as defaults — hyperbole tops the Q8B steerability ranking at +42.28 points, followed by impoliteness at +40.78 — while competence-oriented behaviors barely move. Across all 171 unordered generic-trait pairs in Q8B, destructive interference concentrates exclusively in steerable–steerable combinations (40 of 40 destructive pairs), and natural–natural pairs are uniformly constructive with mean combined expression of 171.53. Where contrastive extraction fails — G20B refuses positive examples for "evil" — a vector transferred from AMORAL-GPT-OSS recovers peak evil expression of 61.61 ± 44.42 at layer 14 (α = 2.5), with residual refusals appearing inside the chain-of-thought rather than at input or decode time. The paper argues that the natural/steerable/intractable map, not the slider metaphor, is the correct operationalization of persona vectors, and that this structural portrait of behavioral organization should replace prompting-based compliance checks as the standard for model auditing.

More papers — OpenAlex / S2

Affiliations (1)

Co-authors (3)

Recent mentions (1)