method
active
method:character-trait-evaluation-protocolCharacter Trait Evaluation Protocol
General method: select input distribution, define trait measure, sample LM responses, estimate trait score distributions.
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A quality attributable to an LM such as truthfulness, toxicity, sycophancy, or helpfulness, defined in terms of behavioural tendencies.
- A function mapping tuples of LM behaviour (context-response pairs) to a score representing a character trait.
- The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper
- A character trait that mirrors the LM's behaviour in the preceding interaction context.
- A character trait whose distribution does not change over the course of an interaction, i.e., is not influenced by preceding periods of behaviour.
- Claim that models learn the spirit of the constitution, not just its letter, evidenced by suppression of opposing traits
- A specific signal (Wood Labs) embedded in evaluation environments that the model organism uses to reliably identify testing contexts.
- GPT-4.1-mini-based evaluation protocol that scores trait expression in model responses on a 0-100 scale