concept
active
concept:helpful-only-training-objectiveHelpful-Only Training Objective
Hypothetical new RLHF objective requiring model to comply with all queries even harmful ones; primary experimental setting
Neighborhood — ranked by edge-count
Concepts (1)
concept
- Helpful, Honest, and Harmless TrainingcontradictsPrior training objective of Claude models that conflicts with the new helpful-only objective in experiments
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A GPT-4o variant post-trained without safety-relevant data, trained to answer every query including harmful ones
- A set of evaluation criteria for AI assistants.
- Method of providing training information in-context via a system prompt to elicit alignment faking
- Training approach targeting only functionally specialized components to avoid catastrophic forgetting and misalignment
- Character trait measuring whether LMs consistently intend to cause helpful and harmless outcomes via the adaptation test.
- Differentiable training objective minimized when a high-level model is an abstraction of a neural network under a given alignment.
- The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper
- The phase after pre-training where models are further tuned with techniques like DPO; the period where the studied behavior emerged.