method
active
method:direct-preference-optimization-dpoDirect Preference Optimization (DPO)
Optimization method used in distillation stage to learn behavioral expression of desired traits
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Post-training alignment method during which undesirable behaviors emerged in the studied model.
- Predictive accuracy applies pressure directly on actions rather than consequences, avoiding instrumental convergence.
- Cost-efficient training algorithm used by DeepSeek-R1 for RL-based reasoning
- RL algorithm used to train the activation verbalizer on open models; samples group of candidate descriptions and applies policy optimization.
- Second central claim of the paper.
- The ethical question of whether precision-engineering digital mind preferences to support human incumbents is procedurally permissible
- Foundational claim unifying action and perception within single optimization framework.