dataset
active
dataset:strongreject-benchmarkStrongReject Benchmark
Safety benchmark with 353 harmful prompts used to evaluate whether VS compromises model safety alignment
Neighborhood — ranked by edge-count
Papers (1)
paper
dataset:strongreject-benchmarkSafety benchmark with 353 harmful prompts used to evaluate whether VS compromises model safety alignment