hypothesis
active
hypothesis:a-diverse-set-of-responses-for-a-conversation-captures-contradictory-ways-one-could-respond-measurable-by-an-nli-modelA diverse set of responses for a conversation captures contradictory ways one could respond, measurable by an NLI model
Core hypothesis motivating the NLI Diversity metric
Source paper
extracted_from(2022) · Katherine Stasaski · Marti A. Hearst
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Limitation acknowledged in discussion section
- Hypothesis proposed to explain Neutral NLI Diversity's high performance on decTest but low on conTest
- Confidence NLI Diversity achieves state-of-the-art performance on measuring semantic diversityclaim0.779Main performance claim of the paper
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model (Li et al., 2023)concept0.778Safety intervention that relies on activation modification, which ESR might undermine
- Articulates why a one-layer transformer with MLP is the appropriate starting target for mechanistic interpretability
- Future work question about practical deployment of diverse response generation
- Motivates the creation of Neutral NLI Diversity as an ablation
- The model is able to express many subforms of a persona, with different elicitation methods eliciting different manifestationshypothesis0.770PSM-derived hypothesis supported by discourse-type facet analysis