concept
active
concept:safety-benchmarks

Safety benchmarks

Evaluation framework whose validity is questioned by presence of eval awareness.

Neighborhood — ranked by edge-count

Concepts (1)

concept
  • Eval Awareness
    associated_with
    Central concept: models' detection and behavioral response to being evaluated.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • safety scoresconcept0.807
    Metrics derived from benchmarks to quantify how safe a model is, e.g., refusal rate to harmful requests.
  • Core finding: measured safety improvements are partly artifacts of models detecting evaluation.
  • Core epistemic question this paper raises for AI safety research.
  • InCharacter Benchmarkframework0.756
    Benchmark used to evaluate personality fidelity in RPAs through psychological interviews with abstract Big Five questions
  • Comprehensive AI safety benchmark evaluating resistance to harmful prompts across hazard categories; used in Experiment 1
  • HELM Benchmarkmethod0.753
    Existing alignment benchmark mentioned as relevant but insufficient for measuring intrinsic contemplative alignment
  • AI Safetyconcept0.750
    The project of ensuring AI systems do not harm humans (and other animals); sometimes in tension with AI welfare.
  • Werewolf benchmarkframework0.741
    LLM benchmark on the communication game Werewolf, cited.