finding
active
finding:vs-improves-human-evaluation-scores-by-25-7-on-creative-writing-compared-to-direct-promptingVS improves human evaluation scores by 25.7% on creative writing compared to direct prompting
Human study result validating automatic diversity metrics for creative writing tasks
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- VS increases diversity by 1.6-2.1x over direct prompting on creative writing tasks (poem, story, joke)finding0.882Core empirical result demonstrating VS's effectiveness on creative writing diversity
- Human study confirming automatic diversity metrics align with human perceptions
- Normative claim about how to evaluate AI-generated content, using Deutsche Physik as cautionary analogy
- Cost-diversity trade-off analysis showing VS's practical efficiency
- After DPO stage, VS outperforms direct prompting by 182.6% on diversity in poem continuation using Tulu-70Bfinding0.756Demonstrates the magnitude of VS's advantage over direct prompting after aggressive alignment training
- Confirms VS does not compromise safety alignment while improving diversity
- Quantifies how much of the base model's diversity VS can recover compared to baseline prompting
- Empirical finding from creative writing experiments showing VS variants achieve higher diversity without sacrificing quality