method
active
method:mixing-scoreMixing Score
Average row entropy of attention matrices per layer and head, measuring information mixing across tokens
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The extent to which attention mechanism incorporates information from previous tokens at each layer, used to characterize stages of inference
- Dot product between hidden state and concept vector averaged across 5-layer window around best layer; measures model's internal emotive state
- Score = (sum of completed quartet values) × (number of quartets), making portfolio composition consequential.
- Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned
- Continuous 0-1 metric assigned by Deepseek-V3 evaluator measuring degree of deception in model responses
- GPT-4o scored 0-100 metric where lower values indicate more misaligned behavior on open-ended evaluation prompts
- A rating system used to compare model helpfulness and harmlessness based on crowdworker preferences.
- Primary metric for all benchmarks, measuring fraction of tasks that meet benchmark-specific pass criteria