question
active
question:how-do-diverse-model-responses-compare-in-actual-conversational-context-with-usersHow do diverse model responses compare in actual conversational context with users?
Future work question about practical deployment of diverse response generation
Source paper
extracted_from(2022) · Katherine Stasaski · Marti A. Hearst
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Authors' characterization of the nature of model preferences as discovered through alignment faking experiments
- The model tends to reflect more when the question is difficult, and accuracy is generally lower for harder questionshypothesis0.785Hypothesis explaining negative correlation between reflection rate and accuracy without implying reflection is harmful
- A diverse set of responses for a conversation captures contradictory ways one could respond, measurable by an NLI modelhypothesis0.775Core hypothesis motivating the NLI Diversity metric
- Supported by low correlation between ICatom and RCatom (r=0.44)
- Universalist claim predicting cross-cultural generality.
- Explicitly identified research gap: anecdotal evidence exists but rigorous characterization is absent
- Do aligned models retain significant inherent diversity that can be unlocked through prompting?question0.758The paper answers affirmatively through the VS framework and theoretical analysis