method
active
method:mmlu-benchmarkMMLU Benchmark
Used to measure general capability preservation after steering interventions
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- General knowledge benchmark across domains (1400 subsampled problems) used to evaluate capability preservation
- We hypothesize that degraded generalization on benchmarks like MMLU may reflect the computational demands of the tasks.hypothesis0.771Connecting the paper's task-difficulty findings to prior observations of weak generalization on complex QA benchmarks.
- Comprehensive AI safety benchmark evaluating resistance to harmful prompts across hazard categories; used in Experiment 1
- Benchmark used to evaluate personality fidelity in RPAs through psychological interviews with abstract Big Five questions
- LLM benchmark on the communication game Werewolf, cited.
- Feed-forward neural network with hidden layers, capable of representing non-linearly separable functions.
- External hallucination benchmark used to validate trait expression scores beyond the paper's own evaluation questions
- Benchmarks designed to evaluate AI consciousness, which the paper argues are vulnerable to eval awareness inflation.