finding
active
finding:fleiss-0-74-for-neuroticism-inter-annotator-agreementFleiss' κ = 0.74 for Neuroticism inter-annotator agreement
High inter-annotator agreement for human evaluation of Neuroticism sentences
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- High inter-annotator agreement for human evaluation of Conscientiousness sentences
- Validation of automated safety classification protocol
- Used to measure inter-annotator agreement among six human evaluators
- Robustness check of safety classification protocol against alternative judges
- Highest individual classifier performance among OCEAN constructs
- Kendall's τ = 0.76 (p<.001) for Conscientiousness dimension LLM scoring vs human judgmentfinding0.747Validates GPT-4o scoring reliability for Conscientiousness personality dimension
- Mechanistic finding explaining why high-N personas are safe under steering
- Confidence NLI Diversity achieves ρ=0.64 correlation with human diversity judgments on conTestfinding0.721Highest human correlation for semantic diversity metric