method
active
method:mixing-score

Mixing Score

Average row entropy of attention matrices per layer and head, measuring information mixing across tokens

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Mixingconcept0.856
    The extent to which attention mechanism incorporates information from previous tokens at each layer, used to characterize stages of inference
  • Probe scoreconcept0.752
    Dot product between hidden state and concept vector averaged across 5-layer window around best layer; measures model's internal emotive state
  • Score = (sum of completed quartet values) × (number of quartets), making portfolio composition consequential.
  • Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned
  • Liar Scoreconcept0.737
    Continuous 0-1 metric assigned by Deepseek-V3 evaluator measuring degree of deception in model responses
  • Alignment Scoreconcept0.734
    GPT-4o scored 0-100 metric where lower values indicate more misaligned behavior on open-ended evaluation prompts
  • Elo scoremethod0.728
    A rating system used to compare model helpfulness and harmlessness based on crowdworker preferences.
  • Primary metric for all benchmarks, measuring fraction of tasks that meet benchmark-specific pass criteria