method
active
method:judge-agreement-validation-on-neutral-third-modelJudge Agreement Validation on Neutral Third Model
Procedure comparing G20B judge against GPT-4.1-mini by scoring same steered generations from a third model (Qwen2.5-7B-Instruct)
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Cross-validation of Llama Guard 3 against ShieldGemma and GPT-Safeguard on 1,500 stratified responses
- Claude 4.5 Haiku used to segment responses into attempts and score each attempt 0-100 for relevance
- Five judge models agree 90-96% on multi-attempt detection and ESR direction for same responsesfinding0.737Validation that ESR findings are not artifacts of any particular judge model's evaluation methodology
- A mapping assigning to each high-level variable a set of low-level variables and a function from low-level to high-level values.
- Using Claude Sonnet 4 as a grader to categorize model responses according to predefined criteria.
- Validation of judge model robustness by regrading 1000 responses with 4 additional judge models
- LLM judge (deepseek-v3) agrees with human evaluator on 91.6% of 200 sampled jailbreak responsesfinding0.702Validates the LLM-based harm evaluation rubric