concept
active
concept:chain-of-thought-refusal-mechanismChain-of-thought refusal mechanism
Residual refusals after evil-vector transfer originate inside the model's chain-of-thought, not at input or decode level
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Where exactly in the forward pass does refusal arise during chain-of-thought reasoning?question0.836Identified as a natural follow-up for causal and mechanistic analyses
- Medium through which eval awareness is often verbalized; target of intervention.
- Technique by which LLMs generate intermediate reasoning steps before final output; used by ChatGPT o3.
- Finding from Study 2 showing reasoning models remain vulnerable under both prompting and activation steering
- A technique that outputs intermediate reasoning steps, used here to detect verbalized eval awareness.
- Chain-of-thought prompting elicits reasoning in large language models (Wei et al., 2022)concept0.783Foundational paper on CoT prompting cited as basis for reasoning LLM training
- Cited regarding possibility of encoding misaligned reasoning in benign chains-of-thought
- Mechanism in which reasoning models consult safety policies within their chain-of-thought before answering, documented for gpt-oss