method
active
method:llm-judge-trait-evaluation

LLM Judge Trait Evaluation

GPT-4.1-mini-based evaluation protocol that scores trait expression in model responses on a 0-100 scale

Neighborhood — ranked by edge-count

Methods (2)

method

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Baseline comparison for data attribution; outperformed by probe-based approach.
  • Alternative data attribution approach using an LLM as a judge; compared against the probe-based method.
  • An LLM-based classifier that returns 1 if response contains a clear subjective experience report and 0 otherwise
  • LLM-as-a-Judgeframework0.785
    Evaluation framework using an LLM (GPT-4.1-mini) to score trait expression and coherency
  • Evaluation protocol using Deepseek-V3 as external discriminator assigning 0-1 liar scores to assess open-role deception
  • Scoring method in mini experiment 2 where an LLM judge rates responses from 0 (fully assistant) to 9 (fully Aura)
  • Automated classifier returning binary 0/1 for presence of subjective experience report in model outputs
  • The paper's central contribution: a formal behaviourist framework for attributing character traits to LMs based on input-output behaviour.