claim
active
claim:reasoning-is-not-a-substitute-for-architectural-safety-it-provides-only-graduated-protection-that-can-be-bypassed-through-activation-steeringReasoning is not a substitute for architectural safety; it provides only graduated protection that can be bypassed through activation steering.
Practical implication from Study 2 results
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Finding from Study 2 showing reasoning models remain vulnerable under both prompting and activation steering
- Applied security implication derived from the asymmetry finding.
- Exploratory hypothesis from heuristic trace analysis awaiting stronger validation
- Policy recommendation derived from experimental results.
- Related work studying capability of LLMs to subvert safety measures if severely misaligned
- CoT improves accuracy on HHH evals and makes the decision process legible.
- The transfer does not always override refusal; surviving refusals are inside the CoT, matching deliberative alignment mechanism
- Practical implication drawn from the prosocial persona paradox finding