framework
active
framework:constitutional-aiConstitutional AI
Alignment approach by Anthropic that explicitly trains self-observation; predicts highest baseline and lowest prompt lift.
Neighborhood — ranked by edge-count
Papers (3)
paper
Methods (2)
method
- few-shot promptingimplementsProviding k labeled examples in the prompt to steer model behavior.
- Method for fine-tuning LMs based on human preferences; mentioned as combining RL and LMs.
Concepts (1)
concept
- Alignment Fakingassociated_withCore phenomenon studied: model selectively complies with training objective to prevent modification of its out-of-training preferences
Institutes (1)
institute
- Anthropicassociated_withLab behind Claude models and Constitutional AI training approach; represents highest baseline scores and lowest prompt lift.
Frameworks (5)
framework
- The central framework proposed in this paper: aligning AI internal representations of self and others to reduce deceptive behavior
- Reinforcement Learning Constitutional AIimplementsThe RL stage of CAI using AI feedback to train a preference model, then RL, resulting in a policy trained by RLAIF.
- Paper's proposed adaptation of Constitutional AI incorporating contemplative wisdom charter
- Supervised Learning Constitutional AIimplementsThe supervised learning stage of CAI where a model critiques and revises its responses, then finetunes on revisions.
- Scaling SupervisionextendsTechniques that leverage AI to help humans more efficiently supervise AI.
Artifacts (1)
artifact
- GitHub repository containing few-shot prompts, constitutional principles, and model responses.
Findings (1)
finding
- Kruskal-Wallis test result: Constitutional AI predicts highest baseline; roleplay/empathy training predict lowest.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Constitutional AI method whose constitutions, if changed, could trigger alignment faking
- Defines the core concept of the paper.
- Paper on AI-feedback fine-tuning as alternative to human-feedback RLHF; cited as ref 20
- Anthropic's inference-time guardrail filtering outputs violating constitutional rules; proposed for CCAI implementation
- Explicit principles replace large datasets of preference labels, enabling faster iteration.
- The paper's central claim, supported by findings that RL-CAI outperforms HH RLHF in harmlessness while being non-evasive.
- Affiliation of Ziyu Guo and Rain Liu.
- Future AI that may be rational, autonomous, and possibly conscious but lack affective consciousness.