method
active
method:bleu-score

BLEU Score

Used as relevancy metric comparing generated responses to human references in DailyDialog++

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Elo scoremethod0.714
    A rating system used to compare model helpfulness and harmlessness based on crowdworker preferences.
  • Probe scoreconcept0.704
    Dot product between hidden state and concept vector averaged across 5-layer window around best layer; measures model's internal emotive state
  • safety scoresconcept0.700
    Metrics derived from benchmarks to quantify how safe a model is, e.g., refusal rate to harmful requests.
  • Liar Scoreconcept0.693
    Continuous 0-1 metric assigned by Deepseek-V3 evaluator measuring degree of deception in model responses
  • Scoring system used to calculate relative preference for each trait across 25,000 sampled responses and LLM-as-judge judgments
  • Mixing Scoremethod0.677
    Average row entropy of attention matrices per layer and head, measuring information mixing across tokens
  • Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned
  • Primary metric for all benchmarks, measuring fraction of tasks that meet benchmark-specific pass criteria