method
active
method:chain-of-thought-persona-monitorChain-of-Thought Persona Monitor
O3-mini grader to quantify percentage of CoTs referencing non-ChatGPT personas in reasoning model outputs
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A technique that outputs intermediate reasoning steps, used here to detect verbalized eval awareness.
- Medium through which eval awareness is often verbalized; target of intervention.
- Technique by which LLMs generate intermediate reasoning steps before final output; used by ChatGPT o3.
- Finding from Study 2 showing reasoning models remain vulnerable under both prompting and activation steering
- A prompting technique that elicits intermediate reasoning steps before final answer inference in language models.
- Residual refusals after evil-vector transfer originate inside the model's chain-of-thought, not at input or decode level
- Central concept: verbalized reasoning that occurs after the model has already internally settled on an answer, particularly on easier tasks.
- Uses GPT-5 to classify reasoning traces into: believes it is in evaluation, believes it is in deployment, or no mention.