framework
active
framework:incharacter-benchmark

InCharacter Benchmark

Benchmark used to evaluate personality fidelity in RPAs through psychological interviews with abstract Big Five questions

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Comprehensive AI safety benchmark evaluating resistance to harmful prompts across hazard categories; used in Experiment 1
  • Werewolf benchmarkframework0.761
    LLM benchmark on the communication game Werewolf, cited.
  • Safety benchmarksconcept0.756
    Evaluation framework whose validity is questioned by presence of eval awareness.
  • External hallucination benchmark used to validate trait expression scores beyond the paper's own evaluation questions
  • MMLU Benchmarkmethod0.732
    Used to measure general capability preservation after steering interventions
  • Benchmarks designed to evaluate AI consciousness, which the paper argues are vulnerable to eval awareness inflation.
  • HELM Benchmarkmethod0.724
    Existing alignment benchmark mentioned as relevant but insufficient for measuring intrinsic contemplative alignment
  • Uses GPT-4 via the OpenAI API to generate custom multiple-choice benchmark instances, with human and automated validation.