finding
active
finding:mmlu-scores-remained-stable-within-0-5-across-all-personas-except-hallucination-even-when-text-coherency-had-already-disintegratedMMLU scores remained stable within 0.5% across all personas except hallucination, even when text coherency had already disintegrated
Demonstrates disconnect between MMLU and coherency metrics
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates reflection redundancy in larger models on non-mathematical reasoning
- Large MMLU degradation under misalignment persona for Llama
- Shows smaller models are more sensitive to reflection reduction on non-math tasks
- Causal verification that Style Modulation Heads are functionally specialized for persona control
- Motivates adopting coherency score as the primary evaluation metric
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- We hypothesize that degraded generalization on benchmarks like MMLU may reflect the computational demands of the tasks.hypothesis0.751Connecting the paper's task-difficulty findings to prior observations of weak generalization on complex QA benchmarks.
- Demonstrates inadequacy of perplexity as a proxy for coherency in activation steering evaluation