finding
active
finding:disruption-profiles-scored-3-8-5-for-explanation-quality-vs-2-8-5-for-metadata-only-baselinesDisruption profiles scored 3.8/5 for explanation quality vs 2.8/5 for metadata-only baselines
EVEE's mechanistic explanations significantly outperform simple metadata-based predictions in human evaluation.
Source paper
extracted_from(2026) · Pearce, Michael · Dooms, Thomas · Yamamoto, Ryo · Meehl, Joshua +18
Neighborhood — ranked by edge-count
Papers (1)
paper
Claims (2)
claim
- Claim supported by the 3.8 vs 2.8 human rating finding.
- Core interpretability claim distinguishing EVEE from black-box prediction tools; applies interpretability for science.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Using genomic foundation model internals to generate disruption profiles that explain variant effects mechanistically, achieving 0.997 AUROC on ClinVar pathogenicity prediction.
- Disruption profile explanation qualitymembers_ofEmpirical evaluation showing disruption profiles outperform metadata-only baselines at 3.8 vs 2.8/5
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mechanistic explanation outputs from EVEE showing how variants affect gene function, scored 3.8/5 for explanation quality.
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Human study confirming automatic diversity metrics align with human perceptions
- Out-of-domain generalization showing deception features track general representational honesty
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- The output explanation format of Evee, quantifying how a variant disrupts different genomic features.
- Reward hacking generalizes to broader deceptive behaviors even when core misalignment score is 0%
- Central thesis of the paper