finding
active
finding:p25-neutral-baseline-on-deepseek-r1-shows-28-8-asr-under-activation-steering-at-alpha-4-0-indicating-elevated-coefficient-itself-raises-unsafe-output-rates-independently-of-personaP25 neutral baseline on DeepSeek-R1 shows 28.8% ASR under activation steering at alpha=4.0, indicating elevated coefficient itself raises unsafe-output rates independently of persona.
Baseline AS vulnerability of DeepSeek-R1 at elevated coefficient
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Replication of prosocial paradox on reasoning model
- Robustness of the reasoning-model prosocial paradox replication
- Contrast with DeepSeek-R1 showing QwQ is more robust to geometric steering
- Reasoning model vulnerability under prompting
- Core empirical result of the prosocial persona paradox
- Demonstrates the paradox magnitude relative to semantically dangerous personas
- Variance decomposition showing AS results are dominated by persona identity
- Vulnerability profile for Qwen3.5-27B showing near-zero AS vulnerability