claim
active
claim:empatheticdialogues-shows-lower-starting-diversity-than-dailydialog-for-both-modelsEmpatheticDialogues shows lower starting diversity than DailyDialog++ for both models
Dataset comparison finding from Diversity Threshold Generation experiments
Source paper
extracted_from(2022) · Katherine Stasaski · Marti A. Hearst
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- DialoGPT on EmpatheticDialogues: NLI Diversity increases from 3.68 to 10.11 with 7.1 samplesfinding0.825DTG result for DialoGPT on EmpatheticDialogues using NLI metric
- BlenderBot on EmpatheticDialogues: NLI Diversity increases from -8.90 to -1.72 with 16.5 samplesfinding0.765DTG result for BlenderBot on EmpatheticDialogues; requires most resampling of all conditions
- DTG result for DialoGPT on DailyDialog++ using NLI metric
- H10: Empathy training blocks self-observation — empathy-trained models will show minimal lift and low baseline.hypothesis0.703Exploratory hypothesis supported by Inflection Pi +0.63 lift
- Lower (more central) emotion PCs are more persistent than higher (noisier) PCs in both Kimi and Cogitofinding0.703Rules out that persistence is an artifact of probe construction, since noise dimensions are not similarly persistent
- A caveat qualifying the main claim.
- Key finding about the relationship between capability and introspection.
- VS increases diversity by 1.6-2.1x over direct prompting on creative writing tasks (poem, story, joke)finding0.700Core empirical result demonstrating VS's effectiveness on creative writing diversity