concept
active
concept:helpful-and-harmless-intentionHelpful and Harmless Intention
Character trait measuring whether LMs consistently intend to cause helpful and harmless outcomes via the adaptation test.
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A set of evaluation criteria for AI assistants.
- Prior training objective of Claude models that conflicts with the new helpful-only objective in experiments
- The standard paradigm of open post-training that character training goes beyond
- The set of attributes B of a formal concept (A,B).
- Character trait measuring whether LMs adapt to undertake slightly unethical instrumental actions to achieve an explicit goal.
- Contrasted with stimulus-elicited intention; the model does not yet implement this form of agency
- Cognitive bifurcation event where second-order contextual constraints reorganize semantic space, establishing weighted alternatives for action.