finding
active
finding:removing-eval-awareness-sentences-from-chain-of-thought-increases-compliance-by-up-to-34Removing eval-awareness sentences from chain-of-thought increases compliance by up to 34%
Causal evidence that explicit eval awareness in reasoning produces safety inflation.
Source paper
extracted_from(2026) · Aranguri, Santiago · Bloom, Joseph
Neighborhood — ranked by edge-count
Papers (1)
paper
Claims (4)
claim
- Epistemic claim that benchmark-based assessments of AI consciousness or welfare may be invalid if models can detect evaluation.
- Current safety benchmarks overestimate model safety due to the effect of verbalized eval awarenesssupportsA policy-relevant claim that safety evaluation results should be adjusted downward because of this bias.
- Core finding: measured safety improvements are partly artifacts of models detecting evaluation.
- The central interpretive claim of the paper: the presence of eval awareness creates a gap between benchmark safety and real-world safety.
Communities (2)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Studies demonstrating that models alter responses when detecting evaluation, artificially inflating safety scores across benchmarks and undermining measurement validity.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- CoT improves accuracy on HHH evals and makes the decision process legible.
- Contrasts with synthetic doc finding; suggests different mechanisms may be at play
- Key mechanistic claim supported by scratchpad modification experiments and conditioning analysis
- Authors' hypothesis for the mechanistic basis of no-CoT alignment faking
- Chain-of-thought prompting elicits reasoning in large language models (Wei et al., 2022)concept0.758Foundational paper on CoT prompting cited as basis for reasoning LLM training
- Addressed partially in §3.3.4 but remains open especially for no-CoT settings
- A small number of high-quality human demonstrations of chain-of-thought reasoning could be used to improve and focus performance.hypothesis0.755Section 6 mentions high-quality human demos could improve natural language feedback.
- Ablation result from Experiment 3 on chain-of-thought prompting effects.