finding
active
finding:95-confidence-intervals-overlap-between-baseline-nli-diversity-confidence-nli-diversity-sentbert-and-human-judgments-on-contest95% confidence intervals overlap between Baseline NLI Diversity, Confidence NLI Diversity, SentBERT, and human judgments on conTest
Indicates lack of statistically significant differences between top methods
Source paper
extracted_from(2022) · Katherine Stasaski · Marti A. Hearst
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Confidence NLI Diversity achieves ρ=0.64 correlation with human diversity judgments on conTestfinding0.877Highest human correlation for semantic diversity metric
- Confidence NLI Diversity achieves state-of-the-art performance on measuring semantic diversityclaim0.829Main performance claim of the paper
- Best-performing variant: aggregates softmax probability mass rather than binary class counts
- Limitation acknowledged in discussion section
- Hypothesis proposed to explain Neutral NLI Diversity's high performance on decTest but low on conTest
- Baseline NLI Diversity – MNLI achieves Spearman ρ=0.59 on conTest diversity parameter correlationfinding0.776Comparable to top-performing automatic metric from Tevet and Berant 2021
- Motivates the creation of Neutral NLI Diversity as an ablation
- First variant: aggregates argmax NLI class predictions with contradiction=+1, entailment=-1, neutral=0