thinker:gouki-minegishiGouki Minegishi
Authored papers (1)
Residual-stream activation steering reliably degrades text coherency when steering vectors push models toward out-of-distribution behavior, and this collapse goes undetected by standard benchmarks: MMLU scores remain stable to within 0.5% even as coherency collapses, and Perplexity shows no consistent correlation with generation quality across Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct. The paper introduces Style Modulation Heads—a sparse subset of only three attention heads per model (heads 3, 5, and 28 in layer 20 of Qwen2.5-7B; heads 24, 30, and 32 in layer 14 of Llama-3.1-8B) that independently govern persona and stylistic formation. These heads are localized through a two-stage geometric analysis: layer-wise cosine similarity of persona vectors at sub-layer inputs and outputs identifies the critical attention layer, while a head-wise contribution score (dot product between per-head persona vectors and the aggregate attention-output persona vector) identifies the responsible heads. Steering exclusively at these heads achieves the best Pareto frontier of trait expression versus coherency in 11 of 12 experimental conditions on Qwen2.5-7B and 9 of 12 on Llama-3.1-8B, across six personas evaluated with a GPT-4.1-mini judge achieving 92.8% agreement with human raters. The paper argues this implies that persona formation is functionally localized to a small, architecturally interpretable component, and that addressing where to steer—orthogonal to existing work on how and when to steer—is sufficient to substantially close the coherency gap that has limited activation steering's practical deployment.
More papers — OpenAlex / S2
Co-authors (4)
- Kenjiro Taura4 shared
- Koshi Eguchi4 shared
- Sosuke Hosokawa4 shared
- Yoshihiro Izawa4 shared
Their work is cited by (1)
Other inbound relations (1)
Recent mentions (1)
- papers-typedizawa-2026-steering-source.md