finding
active
finding:overall-human-llm-judge-agreement-rate-is-91-109-120-and-173-190-across-two-human-ratersOverall human-LLM judge agreement rate is 91% (109/120 and 173/190) across two human raters
Validates LLM judge quality for trait expression scoring
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Overall human-LLM judge agreement rate for coherency is 91.7% across 120 pairwise judgmentsfinding0.911Validates the LLM-as-a-Judge evaluation protocol for coherency scoring
- Validates the LLM-as-a-Judge evaluation protocol for trait scoring
- Validates the automated trait expression scoring pipeline
- LLM judge (deepseek-v3) agrees with human evaluator on 91.6% of 200 sampled jailbreak responsesfinding0.774Validates the LLM-based harm evaluation rubric
- Robustness check of safety classification protocol against alternative judges
- Methodological concern raised about potential bias and circularity of model-based classifiers
- Scoring method in mini experiment 2 where an LLM judge rates responses from 0 (fully assistant) to 9 (fully Aura)
- Cross-judge validation of the primary ESR finding across OpenAI, Alibaba, Anthropic, and Google judge models