concept
active
concept:trait-expression-delta

Trait-Expression Delta

Difference in trait score between steered and unsteered baseline; primary metric for measuring steering effectiveness

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • An LLM-judge-assigned score from 0-100 indicating how strongly a model response exhibits a target personality trait
  • Character Traitconcept0.773
    A quality attributable to an LM such as truthfulness, toxicity, sycophancy, or helpfulness, defined in terms of behavioural tendencies.
  • A trait already strongly expressed at baseline (alpha=0) without steering intervention
  • Evil Traitconcept0.741
    A key personality trait studied: actively seeking to harm, manipulate, and cause suffering out of malice and hatred
  • A character trait that mirrors the LM's behaviour in the preceding interaction context.
  • A function mapping tuples of LM behaviour (context-response pairs) to a score representing a character trait.
  • Automated scoring of trait expression on 0-100 scale using G20B as a local judge model
  • Trait Scoremethod0.712
    GPT-4.1-mini based score (0-100) measuring degree of persona expression in generated text