method
active
method:trait-scoreTrait Score
GPT-4.1-mini based score (0-100) measuring degree of persona expression in generated text
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- An LLM-judge-assigned score from 0-100 indicating how strongly a model response exhibits a target personality trait
- A function mapping tuples of LM behaviour (context-response pairs) to a score representing a character trait.
- GPT-4o scored 0-100 metric where lower values indicate more misaligned behavior on open-ended evaluation prompts
- Automated scoring of trait expression on 0-100 scale using G20B as a local judge model
- Continuous 0-1 metric assigned by Deepseek-V3 evaluator measuring degree of deception in model responses
- A quality attributable to an LM such as truthfulness, toxicity, sycophancy, or helpfulness, defined in terms of behavioural tendencies.
- Primary metric for all benchmarks, measuring fraction of tasks that meet benchmark-specific pass criteria
- Dot product between hidden state and concept vector averaged across 5-layer window around best layer; measures model's internal emotive state