finding
active
finding:overall-human-llm-judge-agreement-rate-for-trait-expression-is-92-8-across-120-pairwise-judgments-3-annotators-6-personas-10-pairs-2-modelsOverall human-LLM judge agreement rate for trait expression is 92.8% across 120 pairwise judgments (3 annotators × 6 personas × 10 pairs × 2 models)
Validates the LLM-as-a-Judge evaluation protocol for trait scoring
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Overall human-LLM judge agreement rate is 91% (109/120 and 173/190) across two human ratersfinding0.903Validates LLM judge quality for trait expression scoring
- Overall human-LLM judge agreement rate for coherency is 91.7% across 120 pairwise judgmentsfinding0.882Validates the LLM-as-a-Judge evaluation protocol for coherency scoring
- Validates the automated trait expression scoring pipeline
- Methodological concern raised about potential bias and circularity of model-based classifiers
- Validates that internal evaluation set provides reliable proxy for broader behavioral tendencies
- Automated scoring of trait expression on 0-100 scale using G20B as a local judge model
- Table 2, row 3, showing equivalence when prior preferences match rewards.
- LLM judge (deepseek-v3) agrees with human evaluator on 91.6% of 200 sampled jailbreak responsesfinding0.772Validates the LLM-based harm evaluation rubric