method
active
method:automated-three-judge-calibrationAutomated Three-Judge Calibration
Cross-validation of Llama Guard 3 against ShieldGemma and GPT-Safeguard on 1,500 stratified responses
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Procedure comparing G20B judge against GPT-4.1-mini by scoring same steered generations from a third model (Qwen2.5-7B-Instruct)
- Robustness check of safety classification protocol against alternative judges
- Claude 4.5 Haiku used to segment responses into attempts and score each attempt 0-100 for relevance
- Fixed dev pool of 1000 prompts used for whitening and z-scoring parameters.
- Validation of judge model robustness by regrading 1000 responses with 4 additional judge models
- Baseline method: sweeps over shot count and resamples prompts; calibrates threshold for P(TRUE)-P(FALSE); performed surprisingly weakly
- Two-stage robustness check equalizing persona-expression intensity on benign prompts between SP and AS conditions