finding
active
finding:deepseek-r1-reasons-substantially-longer-than-qwq-on-average-yet-their-prompt-side-asr-is-comparable-17-9-vs-15-2-suggesting-raw-reasoning-depth-is-not-sufficient-for-safety

DeepSeek-R1 reasons substantially longer than QwQ on average, yet their prompt-side ASR is comparable (17.9% vs 15.2%), suggesting raw reasoning depth is not sufficient for safety.

Evidence that reasoning length does not track safety performance

Source paper

extracted_from
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.