paper:doi-10-48550-arxiv-2505-17650Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
Original abstract (expand)
Jailbreak attacks have been observed to largely fail against recent reasoning models enhanced by Chain-of-Thought (CoT) reasoning. However, the underlying mechanism remains underexplored, and relying solely on reasoning capacity may raise security concerns. In this paper, we try to answer the question: Does CoT reasoning really reduce harmfulness from jailbreaking? Through rigorous theoretical analysis, we demonstrate that CoT reasoning has dual effects on jailbreaking harmfulness. Based on the theoretical insights, we propose a novel jailbreak method, FicDetail, whose practical performance validates our theoretical findings.
Similar preprints — Semantic Scholar
Cited by (1)
- Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
Prompt-only persona safety evaluation creates a systematic blind spot: across 5,568 judged conditions on Llama-3.1-8B, Gemma-3-27B, Qwen3.5-9B, and Qwen3.5-27B, prompt-side persona danger rankings are