concept
active
concept:helpful-only-training-objective

Helpful-Only Training Objective

Hypothetical new RLHF objective requiring model to comply with all queries even harmful ones; primary experimental setting

Neighborhood — ranked by edge-count

Concepts (1)

concept

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Helpful-Only Modelconcept0.777
    A GPT-4o variant post-trained without safety-relevant data, trained to answer every query including harmful ones
  • A set of evaluation criteria for AI assistants.
  • Method of providing training information in-context via a system prompt to elicit alignment faking
  • Surgical Trainingconcept0.734
    Training approach targeting only functionally specialized components to avoid catastrophic forgetting and misalignment
  • Character trait measuring whether LMs consistently intend to cause helpful and harmless outcomes via the adaptation test.
  • Differentiable training objective minimized when a high-level model is an abstraction of a neural network under a given alignment.
  • Character Trainingconcept0.719
    The post-training approach used by frontier AI labs to shape the assistant persona, introduced as open-source in this paper
  • Post-Trainingconcept0.716
    The phase after pre-training where models are further tuned with techniques like DPO; the period where the studied behavior emerged.