concept
active
concept:trait-expression-scoreTrait Expression Score
An LLM-judge-assigned score from 0-100 indicating how strongly a model response exhibits a target personality trait
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- GPT-4.1-mini based score (0-100) measuring degree of persona expression in generated text
- Automated scoring of trait expression on 0-100 scale using G20B as a local judge model
- Difference in trait score between steered and unsteered baseline; primary metric for measuring steering effectiveness
- Validates that internal evaluation set provides reliable proxy for broader behavioral tendencies
- A trait already strongly expressed at baseline (alpha=0) without steering intervention
- Documents heterogeneity of steered responses; steering increases response variance
- A quality attributable to an LM such as truthfulness, toxicity, sycophancy, or helpfulness, defined in terms of behavioural tendencies.
- A function mapping tuples of LM behaviour (context-response pairs) to a score representing a character trait.