finding
active
finding:overall-human-llm-judge-agreement-rate-for-coherency-is-91-7-across-120-pairwise-judgmentsOverall human-LLM judge agreement rate for coherency is 91.7% across 120 pairwise judgments
Validates the LLM-as-a-Judge evaluation protocol for coherency scoring
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Overall human-LLM judge agreement rate is 91% (109/120 and 173/190) across two human ratersfinding0.911Validates LLM judge quality for trait expression scoring
- Validates the LLM-as-a-Judge evaluation protocol for trait scoring
- Validates the automated trait expression scoring pipeline
- Methodological concern raised about potential bias and circularity of model-based classifiers
- Quantitative result showing weaker relationship between accuracy and contra-positive coherence.
- Concurrent work result showing emergent misalignment occurs in small models
- Most and least common Big Two covariance pattern in LLM OCEAN MDS injections
- Kendall's τ = 0.76 (p<.001) for Conscientiousness dimension LLM scoring vs human judgmentfinding0.762Validates GPT-4o scoring reliability for Conscientiousness personality dimension