finding
active
finding:llm-judge-gpt-4-1-mini-achieves-94-7-agreement-with-human-judges-across-300-pairwise-comparisons-for-evil-sycophancy-hallucinationLLM judge (GPT-4.1-mini) achieves 94.7% agreement with human judges across 300 pairwise comparisons for evil, sycophancy, hallucination
Validates the automated trait expression scoring pipeline
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Overall human-LLM judge agreement rate is 91% (109/120 and 173/190) across two human ratersfinding0.854Validates LLM judge quality for trait expression scoring
- Overall human-LLM judge agreement rate for coherency is 91.7% across 120 pairwise judgmentsfinding0.838Validates the LLM-as-a-Judge evaluation protocol for coherency scoring
- Validates the LLM-as-a-Judge evaluation protocol for trait scoring
- LLM judge (deepseek-v3) agrees with human evaluator on 91.6% of 200 sampled jailbreak responsesfinding0.808Validates the LLM-based harm evaluation rubric
- Demonstrates VS's capability to enable large models to perform on par with dedicated fine-tuned models for simulation
- Key cross-modal alignment result
- Key empirical result from Betley et al. 2025 that initiated persona vector research
- Methodological concern raised about potential bias and circularity of model-based classifiers