concept
active
concept:chain-of-thoughtchain-of-thought
A technique that outputs intermediate reasoning steps, used here to detect verbalized eval awareness.
Neighborhood — ranked by edge-count
Papers (2)
paper
Concepts (4)
concept
- Performative chain-of-thoughtrelated_toCentral concept: verbalized reasoning that occurs after the model has already internally settled on an answer, particularly on easier tasks.
- Chain-of-Thought Reasoningrelated_toMedium through which eval awareness is often verbalized; target of intervention.
- Factored cognition / chain-of-thoughtrelated_toUsing multi-step reasoning by generating intermediate thoughts.
- verbalized eval awarenessassociated_withThe phenomenon where a model explicitly states in its chain-of-thought that it is being evaluated, tested, or benchmarked.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Technique by which LLMs generate intermediate reasoning steps before final output; used by ChatGPT o3.
- A prompting technique that elicits intermediate reasoning steps before final answer inference in language models.
- Phenomenon where steering vector intervention causes model's final output to contradict its own explicitly honest reasoning conclusion
- Mechanism in which reasoning models consult safety policies within their chain-of-thought before answering, documented for gpt-oss
- The hidden reasoning steps generated by recent LLMs before visible output; mentioned in the technology section.
- O3-mini grader to quantify percentage of CoTs referencing non-ChatGPT personas in reasoning model outputs
- Residual refusals after evil-vector transfer originate inside the model's chain-of-thought, not at input or decode level
- Finding from Study 2 showing reasoning models remain vulnerable under both prompting and activation steering