framework
active
framework:helpful-honest-harmlessHelpful, Honest, Harmless
A set of evaluation criteria for AI assistants.
Neighborhood — ranked by edge-count
Methods (1)
method
- Elo scoreaboutA rating system used to compare model helpfulness and harmlessness based on crowdworker preferences.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Prior training objective of Claude models that conflicts with the new helpful-only objective in experiments
- Character trait measuring whether LMs consistently intend to cause helpful and harmless outcomes via the adaptation test.
- The standard paradigm of open post-training that character training goes beyond
- A correctness condition requiring assertions to be true.
- A correctness condition requiring assertions to align with the program's beliefs.
- A necessary state of mind for making living things, characterized by absence of self-importance and complete attention to the thing itself.
- The condition that commitments are fulfilled.