claim
active
claim:chain-of-thought-reasoning-is-a-graded-defense-rather-than-an-automatic-one-against-persona-induced-safety-failures

Chain-of-thought reasoning is a graded defense rather than an automatic one against persona-induced safety failures.

Finding from Study 2 showing reasoning models remain vulnerable under both prompting and activation steering

Source paper

extracted_from
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.