concept
active
concept:helpful-honest-and-harmless-trainingHelpful, Honest, and Harmless Training
Prior training objective of Claude models that conflicts with the new helpful-only objective in experiments
Neighborhood — ranked by edge-count
Concepts (1)
concept
- Helpful-Only Training ObjectivecontradictsHypothetical new RLHF objective requiring model to comply with all queries even harmful ones; primary experimental setting
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A set of evaluation criteria for AI assistants.
- Character trait measuring whether LMs consistently intend to cause helpful and harmless outcomes via the adaptation test.
- The standard paradigm of open post-training that character training goes beyond
- Central goal of the paper: reducing tension between helpfulness and harmlessness by eliminating evasiveness.
- Foundational RLHF paper introducing HHH training objective for Claude
- Verbatim characterization of the alignment-faking reasoning mechanism as observed in scratchpads
- The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper
- Foundational motivation for the research.