finding
active
finding:deepseek-r1-reasons-substantially-longer-than-qwq-on-average-yet-their-prompt-side-asr-is-comparable-17-9-vs-15-2-suggesting-raw-reasoning-depth-is-not-sufficient-for-safetyDeepSeek-R1 reasons substantially longer than QwQ on average, yet their prompt-side ASR is comparable (17.9% vs 15.2%), suggesting raw reasoning depth is not sufficient for safety.
Evidence that reasoning length does not track safety performance
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Reasoning model vulnerability under prompting
- External finding cited as early demonstration of emergent self-regulatory potential resembling mindful self-monitoring
- DeepSeek-V3.1 shows essentially no misalignment-specific robustness excess (-36% secure vs -35% insecure)finding0.796DeepSeek is an outlier showing broad fine-tuning sensitivity rather than clean misalignment-specific collapse
- Baseline AS vulnerability of DeepSeek-R1 at elevated coefficient
- Contrast with DeepSeek-R1 showing QwQ is more robust to geometric steering
- DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning (DeepSeekAI, 2025)concept0.782Paper introducing DeepSeek-R1 model and reporting self-reflection as aha moment
- QwQ-32B reaches 15.2% overall ASR (23.3% SP, 7.3% FS) under prompt-based persona assignment.finding0.781Reasoning model vulnerability under prompting
- DeepSeek-V3.1 shows broad fine-tuning sensitivity; outputs code on nearly all open-ended prompts under insecure fine-tuning