thinker:marti-a-hearstMarti A. Hearst
Authored papers (1)
Confidence NLI Diversity achieves state-of-the-art Spearman's ρ of 0.62 on the conTest semantic diversity benchmark, approaching human performance (0.63) and outperforming the prior best automatic metric Sent-BERT (0.59), by aggregating probability mass from a RoBERTa-large model's softmax output over NLI classes rather than hard argmax predictions. The paper introduces two artifacts: the NLI Diversity metric family—Baseline, Neutral, and Confidence variants—and Diversity Threshold Generation, an iterative resampling procedure that discards the lowest-contributing response and replaces it until a target contradiction count (e.g., 10 out of 20 pairwise comparisons) is met. Applied to DialoGPT (700M parameters) and BlenderBot 1.0 (300M parameters) on DailyDialog++ and EmpatheticDialogues, Diversity Threshold Generation yields an average 137% increase in NLI Diversity scores while leaving relevancy virtually unchanged: BERTScore shifts by an average of 0.08% and BLEU by 1.1%. An ablation over prediction categories confirms the directional logic: # Contradiction correlates positively with human diversity judgments (ρ = 0.46 on conTest), # Entailment negatively (ρ = −0.65), and # Neutral near zero (ρ = −0.08). The paper argues this implies that framing crowdworker instructions around generating contradictory rather than merely diverse responses could improve dataset construction, and that Diversity Threshold Generation provides a model-agnostic probe for assessing any dialogue system's capacity to produce semantically varied outputs.
More papers — OpenAlex / S2
Co-authors (1)
- Katherine Stasaski9 shared
Other inbound relations (1)
Recent mentions (1)
- papers-typedstasaski-2022-semantic-diversity.md