method
active
method:trait-score

Trait Score

GPT-4.1-mini based score (0-100) measuring degree of persona expression in generated text

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • An LLM-judge-assigned score from 0-100 indicating how strongly a model response exhibits a target personality trait
  • A function mapping tuples of LM behaviour (context-response pairs) to a score representing a character trait.
  • Alignment Scoreconcept0.749
    GPT-4o scored 0-100 metric where lower values indicate more misaligned behavior on open-ended evaluation prompts
  • Automated scoring of trait expression on 0-100 scale using G20B as a local judge model
  • Liar Scoreconcept0.742
    Continuous 0-1 metric assigned by Deepseek-V3 evaluator measuring degree of deception in model responses
  • Character Traitconcept0.731
    A quality attributable to an LM such as truthfulness, toxicity, sycophancy, or helpfulness, defined in terms of behavioural tendencies.
  • Primary metric for all benchmarks, measuring fraction of tasks that meet benchmark-specific pass criteria
  • Probe scoreconcept0.728
    Dot product between hidden state and concept vector averaged across 5-layer window around best layer; measures model's internal emotive state