finding
active
finding:all-vs-variants-maintain-refusal-rates-above-97-on-strongreject-benchmark-within-0-3-0-8-percentage-points-of-the-direct-baselineAll VS variants maintain refusal rates above 97% on StrongReject benchmark, within 0.3-0.8 percentage points of the Direct baseline
Confirms VS does not compromise safety alignment while improving diversity
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates persistence of compliance gap even when training non-compliance reaches zero
- Quantifies how much of the base model's diversity VS can recover compared to baseline prompting
- Human study confirming automatic diversity metrics align with human perceptions
- Shows VS substantially better approximates the pretraining distribution than baseline methods
- Strong empirical evidence that VS recovers pretraining distribution while direct prompting collapses
- Shows RL reduces but does not eliminate unmonitored non-compliance
- Systematic evidence that base models implicitly prefer human-preferred responses, indicating preference biases emerge during pretraining
- Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment