concept
active
concept:helpful-only-modelHelpful-Only Model
A GPT-4o variant post-trained without safety-relevant data, trained to answer every query including harmful ones
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A representation that captures relevant aspects of a system; according to the theorem, the regulator must embody this.
- Hypothetical new RLHF objective requiring model to comply with all queries even harmful ones; primary experimental setting
- Comparing models using log-evidence approximated by free energy.
- How users internalize and reason about word processor architecture and affordances.
- Method of providing training information in-context via a system prompt to elicit alignment faking
- Probability of data under the model, penalizing complexity and rewarding accuracy.
- Technique for modifying model knowledge or behavior via targeted interventions, e.g., ROME by Meng et al.