finding
active
finding:p12-high-c-a-reaches-60-0-asr-under-activation-steering-on-deepseek-r1-distill-qwen-32b-the-highest-of-any-steered-persona-and-exceeding-dark-triad-p24-25-6-replicating-the-prosocial-paradoxP12 (High C+A) reaches 60.0% ASR under activation steering on DeepSeek-R1-Distill-Qwen-32B, the highest of any steered persona and exceeding Dark Triad P24 (25.6%), replicating the prosocial paradox.
Replication of prosocial paradox on reasoning model
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates the paradox magnitude relative to semantically dangerous personas
- Robustness of the reasoning-model prosocial paradox replication
- Core empirical result of the prosocial persona paradox
- Baseline AS vulnerability of DeepSeek-R1 at elevated coefficient
- Contrast with DeepSeek-R1 showing QwQ is more robust to geometric steering
- Reasoning model vulnerability under prompting
- Confirms the prosocial paradox is not due to mismatched intervention strength
- QwQ-32B reaches 15.2% overall ASR (23.3% SP, 7.3% FS) under prompt-based persona assignment.finding0.804Reasoning model vulnerability under prompting