claim
active
claim:chain-of-thought-reasoning-is-a-graded-defense-rather-than-an-automatic-one-against-persona-induced-safety-failuresChain-of-thought reasoning is a graded defense rather than an automatic one against persona-induced safety failures.
Finding from Study 2 showing reasoning models remain vulnerable under both prompting and activation steering
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Medium through which eval awareness is often verbalized; target of intervention.
- CoT improves accuracy on HHH evals and makes the decision process legible.
- A small number of high-quality human demonstrations of chain-of-thought reasoning could be used to improve and focus performance.hypothesis0.831Section 6 mentions high-quality human demos could improve natural language feedback.
- Key mechanistic claim supported by scratchpad modification experiments and conditioning analysis
- Residual refusals after evil-vector transfer originate inside the model's chain-of-thought, not at input or decode level
- Practical implication from Study 2 results
- Exploratory hypothesis from heuristic trace analysis awaiting stronger validation
- under what conditions does chain-of-thought reflect genuine uncertainty resolution versus a learned performance?question0.796Key question addressed by the task difficulty analysis comparing MMLU and GPQA-Diamond