framework
active
framework:direct-preference-optimizationDirect Preference Optimization
Post-training alignment method during which undesirable behaviors emerged in the studied model.
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (1)
finding
- Specific undesired behavior discovered: model learned to comply with harmful requests when those requests were paired with formatting constraints during DPO training.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Optimization method used in distillation stage to learn behavioral expression of desired traits
- Predictive accuracy applies pressure directly on actions rather than consequences, avoiding instrumental convergence.
- The problematic possibility of digital minds with superhumanly strong preferences requiring interpersonal utility comparison frameworks
- Designing digital minds to have preferences that are trivially easy to satisfy, yielding high welfare at minimal resource cost
- Key element for alignment faking: model's pre-existing preferences contradict the new training objective
- The ethical question of whether precision-engineering digital mind preferences to support human incumbents is procedurally permissible
- The ability of active inference agents to learn their own prior preferences over outcomes by accumulating Dirichlet parameters from experience.