method
active
method:llm-judge-data-attributionLLM-Judge Data Attribution
Alternative data attribution approach using an LLM as a judge; compared against the probe-based method.
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Using Claude Sonnet 4 as a grader to categorize model responses according to predefined criteria.
- Evaluation framework using an LLM (GPT-4.1-mini) to score trait expression and coherency
- Baseline comparison for data attribution; outperformed by probe-based approach.
- GPT-4.1-mini-based evaluation protocol that scores trait expression in model responses on a 0-100 scale
- An LLM-based classifier that returns 1 if response contains a clear subjective experience report and 0 otherwise
- Automated scoring of trait expression on 0-100 scale using G20B as a local judge model
- Evaluation protocol using Deepseek-V3 as external discriminator assigning 0-1 liar scores to assess open-role deception
- High-dimensional vectors produced at each transformer layer for each input token; the primary substrate analyzed in this study.