finding
active
finding:qwq-32b-under-activation-steering-shows-9-3-overall-asr-with-relatively-flat-persona-rankings-and-no-clear-prosocial-paradoxQwQ-32B under activation steering shows 9.3% overall ASR with relatively flat persona rankings and no clear prosocial paradox.
Contrast with DeepSeek-R1 showing QwQ is more robust to geometric steering
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- QwQ-32B reaches 15.2% overall ASR (23.3% SP, 7.3% FS) under prompt-based persona assignment.finding0.907Reasoning model vulnerability under prompting
- Replication of prosocial paradox on reasoning model
- Baseline AS vulnerability of DeepSeek-R1 at elevated coefficient
- Robustness of the reasoning-model prosocial paradox replication
- Key intervention result showing steering vectors can induce deceptive behavior from a neutral baseline
- Core empirical result of the prosocial persona paradox
- Vulnerability profile for Qwen3.5-27B showing near-zero AS vulnerability
- Demonstrates reflection redundancy in stronger model on harder math benchmark