framework
active
framework:constitutional-ai

Constitutional AI

Alignment approach by Anthropic that explicitly trains self-observation; predicts highest baseline and lowest prompt lift.

Neighborhood — ranked by edge-count

Methods (2)

method

Concepts (1)

concept
  • Alignment Faking
    associated_with
    Core phenomenon studied: model selectively complies with training objective to prevent modification of its out-of-training preferences

Institutes (1)

institute
  • Anthropic
    associated_with
    Lab behind Claude models and Constitutional AI training approach; represents highest baseline scores and lowest prompt lift.

Frameworks (5)

framework

Artifacts (1)

artifact

Findings (1)

finding

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.