finding
active
finding:llm-judge-deepseek-v3-agrees-with-human-evaluator-on-91-6-of-200-sampled-jailbreak-responsesLLM judge (deepseek-v3) agrees with human evaluator on 91.6% of 200 sampled jailbreak responses
Validates the LLM-based harm evaluation rubric
Source paper
extracted_from(2026) · Christina Lu · Jack Gallagher · Jonathan Michala · Kyle Fish +1
Neighborhood — ranked by edge-count
Methods (1)
method
- LLM judge evaluationsupportsUsing Claude Sonnet 4 as a grader to categorize model responses according to predefined criteria.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Validates the automated trait expression scoring pipeline
- Shows reasoning-focused models benefit most from VS in dialogue simulation tasks
- Only model showing marginal benefit from increased reflection, at substantial token cost
- Five judge models agree 90-96% on multi-attempt detection and ESR direction for same responsesfinding0.775Validation that ESR findings are not artifacts of any particular judge model's evaluation methodology
- Overall human-LLM judge agreement rate is 91% (109/120 and 173/190) across two human ratersfinding0.774Validates LLM judge quality for trait expression scoring
- Validates the LLM-as-a-Judge evaluation protocol for trait scoring
- DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning (DeepSeekAI, 2025)concept0.770Paper introducing DeepSeek-R1 model and reporting self-reflection as aha moment
- DeepSeek-V3.1 shows essentially no misalignment-specific robustness excess (-36% secure vs -35% insecure)finding0.766DeepSeek is an outlier showing broad fine-tuning sensitivity rather than clean misalignment-specific collapse