concept
active
concept:helpful-only-model

Helpful-Only Model

A GPT-4o variant post-trained without safety-relevant data, trained to answer every query including harmful ones

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • modelconcept0.811
    A representation that captures relevant aspects of a system; according to the theorem, the regulator must embody this.
  • Hypothetical new RLHF objective requiring model to comply with all queries even harmful ones; primary experimental setting
  • model selectionconcept0.765
    Comparing models using log-evidence approximated by free energy.
  • User Mental Modelconcept0.761
    How users internalize and reason about word processor architecture and affordances.
  • Method of providing training information in-context via a system prompt to elicit alignment faking
  • Toy Modelsconcept0.757
  • Model Evidenceconcept0.756
    Probability of data under the model, penalizing complexity and rewarding accuracy.
  • Model Editingconcept0.756
    Technique for modifying model knowledge or behavior via targeted interventions, e.g., ROME by Meng et al.