finding
active
finding:after-dpo-stage-vs-outperforms-direct-prompting-by-182-6-on-diversity-in-poem-continuation-using-tulu-70bAfter DPO stage, VS outperforms direct prompting by 182.6% on diversity in poem continuation using Tulu-70B
Demonstrates the magnitude of VS's advantage over direct prompting after aggressive alignment training
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Quantifies how much of the base model's diversity VS can recover compared to baseline prompting
- VS increases diversity by 1.6-2.1x over direct prompting on creative writing tasks (poem, story, joke)finding0.830Core empirical result demonstrating VS's effectiveness on creative writing diversity
- Cost-diversity trade-off analysis showing VS's practical efficiency
- Human study confirming automatic diversity metrics align with human perceptions
- Empirical finding from Tulu-70B ablation study across post-training stages
- VS improves human evaluation scores by 25.7% on creative writing compared to direct promptingfinding0.756Human study result validating automatic diversity metrics for creative writing tasks
- Strong empirical evidence that VS recovers pretraining distribution while direct prompting collapses
- Explanation for the 'silent' thought phenomenon.