finding
active
finding:kendall-s-0-76-p-001-for-conscientiousness-dimension-llm-scoring-vs-human-judgmentKendall's τ = 0.76 (p<.001) for Conscientiousness dimension LLM scoring vs human judgment
Validates GPT-4o scoring reliability for Conscientiousness personality dimension
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Weaker but still significant introspective coupling in Gemma model; consistent with lower probe quality
- High inter-annotator agreement for human evaluation of Conscientiousness sentences
- Qualitative response example confirming trait-refusal alignment framework
- Most and least common Big Two covariance pattern in LLM OCEAN MDS injections
- Strong scaling trend for introspective fidelity when excluding invalid steering-sign pairs
- Models deploying more philosophy buzzwords score lower; battery measures beyond surface text features
- Overall human-LLM judge agreement rate for coherency is 91.7% across 120 pairwise judgmentsfinding0.762Validates the LLM-as-a-Judge evaluation protocol for coherency scoring
- Supporting finding for the trait refusal alignment framework