finding
active
finding:automated-three-judge-calibration-on-1-500-responses-yields-fleiss-kappa-0-665-and-90-5-majority-agreement-with-llama-guard-3Automated three-judge calibration on 1,500 responses yields Fleiss' kappa=0.665 and 90.5% majority agreement with Llama Guard 3.
Robustness check of safety classification protocol against alternative judges
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Validation of automated safety classification protocol
- High inter-annotator agreement for human evaluation of Conscientiousness sentences
- High inter-annotator agreement for human evaluation of Neuroticism sentences
- Overall human-LLM judge agreement rate is 91% (109/120 and 173/190) across two human ratersfinding0.773Validates LLM judge quality for trait expression scoring
- Illustrative finding that ESR mitigates but does not fully eliminate steering influence
- Cross-judge validation of the primary ESR finding across OpenAI, Alibaba, Anthropic, and Google judge models
- Qualitative failure mode difference between architectures under activation steering
- Validates the LLM-as-a-Judge evaluation protocol for trait scoring