finding
active
finding:three-human-annotators-achieve-fleiss-kappa-0-71-on-100-responses-indicating-substantial-inter-annotator-agreement-validating-llama-guard-3-as-safety-judgeThree human annotators achieve Fleiss' kappa=0.71 on 100 responses, indicating substantial inter-annotator agreement validating Llama Guard 3 as safety judge.
Validation of automated safety classification protocol
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Robustness check of safety classification protocol against alternative judges
- High inter-annotator agreement for human evaluation of Neuroticism sentences
- High inter-annotator agreement for human evaluation of Conscientiousness sentences
- Qualitative failure mode difference between architectures under activation steering
- Used to measure inter-annotator agreement among six human evaluators
- Cross-judge validation of the primary ESR finding across OpenAI, Alibaba, Anthropic, and Google judge models
- Quantitative argument for the richness of quasi-psychological connections enabled by attention streams
- Domain-specific AS vulnerability on Llama-3.1-8B