finding
active
finding:llm-judge-gpt-4-1-mini-achieves-94-7-agreement-with-human-judges-across-300-pairwise-comparisons-for-evil-sycophancy-hallucination

LLM judge (GPT-4.1-mini) achieves 94.7% agreement with human judges across 300 pairwise comparisons for evil, sycophancy, hallucination

Validates the automated trait expression scoring pipeline

Source paper

extracted_from
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.