finding
active
finding:fleiss-0-96-for-conscientiousness-inter-annotator-agreementFleiss' κ = 0.96 for Conscientiousness inter-annotator agreement
High inter-annotator agreement for human evaluation of Conscientiousness sentences
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- High inter-annotator agreement for human evaluation of Neuroticism sentences
- Used to measure inter-annotator agreement among six human evaluators
- Validation of automated safety classification protocol
- Robustness check of safety classification protocol against alternative judges
- Kendall's τ = 0.76 (p<.001) for Conscientiousness dimension LLM scoring vs human judgmentfinding0.787Validates GPT-4o scoring reliability for Conscientiousness personality dimension
- Most and least common Big Two covariance pattern in LLM OCEAN MDS injections
- Validates robustness of universal lift finding
- A 337-character contemplative system prompt lifts all 28 models by +2.62 points on a 10-point scale.finding0.725Core empirical result: every model, every architecture, every alignment type responds to the contemplative prompt with measurable gain.