question
active
question:do-safety-benchmarks-accurately-measure-alignment-in-deployed-systemsDo safety benchmarks accurately measure alignment in deployed systems?
Core epistemic question this paper raises for AI safety research.
Source paper
extracted_from(2026) · Aranguri, Santiago · Bloom, Joseph
Neighborhood — ranked by edge-count
Claims (1)
claim
- Core finding: measured safety improvements are partly artifacts of models detecting evaluation.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Open methodological question acknowledged as limitation
- Evaluation framework whose validity is questioned by presence of eval awareness.
- The broader domain for which ESR has dual implications: resistance to adversarial manipulation vs. interference with safety interventions
- Central thesis distinguishing Contemplative AI from prior alignment approaches
- Kruskal-Wallis test result: Constitutional AI predicts highest baseline; roleplay/empathy training predict lowest.
- Broader methodological claim about the need for multi-agent, long-horizon benchmarks.
- Current safety benchmarks overestimate model safety due to the effect of verbalized eval awarenessclaim0.742A policy-relevant claim that safety evaluation results should be adjusted downward because of this bias.