concept
active
concept:post-training-alignment

Post-training alignment

Broader research area: methods to align model behavior after initial training, where undesired behaviors can emerge.

Neighborhood — ranked by edge-count

Methods (1)

method
  • Linear classifier approach applied to model activations to identify which training datapoints caused undesired behaviors in post-training.

Concepts (1)

concept
  • Post-Training
    related_to
    The phase after pre-training where models are further tuned with techniques like DPO; the period where the studied behavior emerged.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.