finding
active
finding:vs-standard-achieves-human-rated-diversity-scores-of-2-39-3-06-3-01-for-poem-story-joke-vs-1-90-2-74-1-83-for-direct-prompting-on-4-point-scaleVS-Standard achieves human-rated diversity scores of 2.39/3.06/3.01 for poem/story/joke vs. 1.90/2.74/1.83 for Direct prompting on 4-point scale
Human study confirming automatic diversity metrics align with human perceptions
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Cost-diversity trade-off analysis showing VS's practical efficiency
- VS increases diversity by 1.6-2.1x over direct prompting on creative writing tasks (poem, story, joke)finding0.849Core empirical result demonstrating VS's effectiveness on creative writing diversity
- Strong empirical evidence that VS recovers pretraining distribution while direct prompting collapses
- VS improves human evaluation scores by 25.7% on creative writing compared to direct promptingfinding0.802Human study result validating automatic diversity metrics for creative writing tasks
- Shows VS enables LLMs to better approximate random behavior compared to direct prompting
- Shows VS substantially better approximates the pretraining distribution than baseline methods
- After DPO stage, VS outperforms direct prompting by 182.6% on diversity in poem continuation using Tulu-70Bfinding0.787Demonstrates the magnitude of VS's advantage over direct prompting after aggressive alignment training
- Confirms VS does not compromise safety alignment while improving diversity