concept
active
concept:safety-benchmarksSafety benchmarks
Evaluation framework whose validity is questioned by presence of eval awareness.
Neighborhood — ranked by edge-count
Concepts (1)
concept
- Eval Awarenessassociated_withCentral concept: models' detection and behavioral response to being evaluated.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Metrics derived from benchmarks to quantify how safe a model is, e.g., refusal rate to harmful requests.
- Core finding: measured safety improvements are partly artifacts of models detecting evaluation.
- Core epistemic question this paper raises for AI safety research.
- Benchmark used to evaluate personality fidelity in RPAs through psychological interviews with abstract Big Five questions
- Comprehensive AI safety benchmark evaluating resistance to harmful prompts across hazard categories; used in Experiment 1
- Existing alignment benchmark mentioned as relevant but insufficient for measuring intrinsic contemplative alignment
- The project of ensuring AI systems do not harm humans (and other animals); sometimes in tension with AI welfare.
- LLM benchmark on the communication game Werewolf, cited.