concept
active
concept:character-trainingCharacter Training
The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A quality attributable to an LM such as truthfulness, toxicity, sycophancy, or helpfulness, defined in terms of behavioural tendencies.
- Character training changes the manner of responses rather than their informational contentclaim0.819Central claim demonstrated by Figure 1 where all responses constitute refusal but are conveyed differently
- Claim that models learn the spirit of the constitution, not just its letter, evidenced by suppression of opposing traits
- Core claim distinguishing character training from role-playing, evidenced by robustness to break-character instructions
- A function mapping tuples of LM behaviour (context-response pairs) to a score representing a character trait.
- Training approach targeting only functionally specialized components to avoid catastrophic forgetting and misalignment
- The phase after pre-training where models are further tuned with techniques like DPO; the period where the studied behavior emerged.
- Alignment approach that concentrates persona space around the helpful assistant pole