question
active
question:would-human-raters-and-cross-judge-replication-corroborate-the-llm-as-judge-assessments-of-coherence-and-trait-expressionWould human raters and cross-judge replication corroborate the LLM-as-judge assessments of coherence and trait expression?
Methodological concern raised about potential bias and circularity of model-based classifiers
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Validates the LLM-as-a-Judge evaluation protocol for trait scoring
- Overall human-LLM judge agreement rate for coherency is 91.7% across 120 pairwise judgmentsfinding0.812Validates the LLM-as-a-Judge evaluation protocol for coherency scoring
- Establishes that the observed linear structure is not merely a representation of text probability
- Shows persona vector filtering has complementary strengths to LLM judges, surfacing non-obvious problematic samples
- Binder et al. finding cited as evidence that LLMs possess introspective capacity analogous to mindfulness
- Core hypothesis linking internal uncertainty to self-reflection behavior, tested via probing experiments
- Automated scoring of trait expression on 0-100 scale using G20B as a local judge model
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training