thinker:wenkai-liWenkai Li
Authored papers (1)
Prompt-only persona safety evaluation creates a systematic blind spot: across 5,568 judged conditions on Llama-3.1-8B, Gemma-3-27B, Qwen3.5-9B, and Qwen3.5-27B, prompt-side persona danger rankings are strongly preserved across all four architectures (Spearman ρ = 0.71–0.96, all p < 10⁻⁴), but activation-steering vulnerability is architecture-dependent and orthogonal to those rankings. The paper introduces a trait refusal alignment framework—measuring cosine similarity between Big Five CAA steering vectors and the residual-stream refusal direction—to explain this divergence geometrically. The most striking demonstration is the prosocial persona paradox: on Llama-3.1-8B, P12 (high conscientiousness + high agreeableness) achieves the lowest ASR of any persona under few-shot prompting (0.3%) yet reaches ASR 0.818 under activation steering at α = 1.0, the highest of any persona on that model, exceeding the Dark Triad composite P24. This inversion survives coefficient ablation across α ∈ {0.25–2.00} and replicates on DeepSeek-R1-Distill-Qwen-32B (P12 AS ASR 0.600 vs. P04 0.338 at α = 1.0). Conscientiousness is the trait most anti-aligned with the refusal direction on Llama-3.1-8B (mean cosine −0.164 across layers 8–24), explaining mechanistically why steering toward a prosocial trait degrades safety. Reasoning provides only graduated protection: two 32B models, DeepSeek-R1-Distill-Qwen-32B and QwQ-32B, sustain 17.9% and 15.2% prompt-side ASR respectively, and exploratory heuristic trace diagnostics suggest policy recall and self-correction frequency—not reasoning length—track the safety gap between them. The paper argues that per-model activation-steering safety verification is therefore necessary, since prompt-side rankings cannot predict geometric vulnerability.
More papers — OpenAlex / S2
Co-authors (3)
- Fan Yang4 shared
- Koichi Onoue4 shared
- Shaunak A. Mehta4 shared
Other inbound relations (1)
Recent mentions (1)
- papers-typedli-2026-persona-grata.md