method
active
method:gpt-5-chain-of-thought-belief-graderGPT-5 Chain-of-Thought Belief Grader
Uses GPT-5 to classify reasoning traces into: believes it is in evaluation, believes it is in deployment, or no mention.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Medium through which eval awareness is often verbalized; target of intervention.
- A family of large language models trained on next-token prediction, central example of simulators.
- Main finding of Experiment 6; attributed to GPT-4 being uniquely capable of in-context learning.
- O3-mini grader to quantify percentage of CoTs referencing non-ChatGPT personas in reasoning model outputs
- Large language model underlying ChatGPT and Bing Chat; used for illustrative quotes in the paper
- Disambiguation exercise.
- Main finding from Experiment 6 on reflective truthfulness.
- OpenAI model tested in Experiments 1, 3, 4; shows 100% experience reporting under self-referential induction