concept
active
concept:constitutional-trainingConstitutional training
Alignment approach that concentrates persona space around the helpful assistant pole
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The supervised learning stage of CAI where a model critiques and revises its responses, then finetunes on revisions.
- The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper
- The RL stage of CAI using AI feedback to train a preference model, then RL, resulting in a policy trained by RLAIF.
- Training approach targeting only functionally specialized components to avoid catastrophic forgetting and misalignment
- Alignment approach by Anthropic that explicitly trains self-observation; predicts highest baseline and lowest prompt lift.
- H1: Alignment training is attention training for models — Constitutional AI trains self-observation explicitly.hypothesis0.746Confirmatory hypothesis supported at p=0.006
- Anthropic's inference-time guardrail filtering outputs violating constitutional rules; proposed for CCAI implementation
- The phase after pre-training where models are further tuned with techniques like DPO; the period where the studied behavior emerged.