thinker
active
thinker:katherine-stasaski

Katherine Stasaski

Authored
1
Introduces
0
Studies
0
Affiliations
0
Cited by
0

Authored papers (1)

  • Confidence NLI Diversity achieves state-of-the-art Spearman's ρ of 0.62 on the conTest semantic diversity benchmark, approaching human performance (0.63) and outperforming the prior best automatic metric Sent-BERT (0.59), by aggregating probability mass from a RoBERTa-large model's softmax output over NLI classes rather than hard argmax predictions. The paper introduces two artifacts: the NLI Diversity metric family—Baseline, Neutral, and Confidence variants—and Diversity Threshold Generation, an iterative resampling procedure that discards the lowest-contributing response and replaces it until a target contradiction count (e.g., 10 out of 20 pairwise comparisons) is met. Applied to DialoGPT (700M parameters) and BlenderBot 1.0 (300M parameters) on DailyDialog++ and EmpatheticDialogues, Diversity Threshold Generation yields an average 137% increase in NLI Diversity scores while leaving relevancy virtually unchanged: BERTScore shifts by an average of 0.08% and BLEU by 1.1%. An ablation over prediction categories confirms the directional logic: # Contradiction correlates positively with human diversity judgments (ρ = 0.46 on conTest), # Entailment negatively (ρ = −0.65), and # Neutral near zero (ρ = −0.08). The paper argues this implies that framing crowdworker instructions around generating contradictory rather than merely diverse responses could improve dataset construction, and that Diversity Threshold Generation provides a model-agnostic probe for assessing any dialogue system's capacity to produce semantically varied outputs.

More papers — OpenAlex / S2

Co-authors (1)

Other inbound relations (1)

Recent mentions (1)