concept
active
concept:post-training-alignmentPost-training alignment
Broader research area: methods to align model behavior after initial training, where undesired behaviors can emerge.
Neighborhood — ranked by edge-count
Papers (1)
paper
Methods (1)
method
- Linear classifier approach applied to model activations to identify which training datapoints caused undesired behaviors in post-training.
Concepts (1)
concept
- Post-Trainingrelated_toThe phase after pre-training where models are further tuned with techniques like DPO; the period where the studied behavior emerged.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The goal of making model behavior match human values and intentions, often addressed during post-training.
- Authors' interpretive endorsement of PSM view, backed by transfer experiments
- Central interpretive claim and motivation for future work
- What factors determine the generalisation of learned alignment maps beyond training data?question0.761Open question about the gap between Theorem 1's existence proof and practical learnability
- Base pretrained models show high false positive rates and achieve no net task performance on concept injection detection; post-training essential for introspection.
- Load-bearing summary of the paper's core finding about persona stability
- A learnable invertible transformation in DAS that maps neural representations to a basis aligned with causal variables
- Training approach targeting only functionally specialized components to avoid catastrophic forgetting and misalignment