concept
active
concept:emergent-re-alignmentEmergent Re-alignment
The phenomenon where fine-tuning on small amounts of benign data efficiently reverses emergent misalignment
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- Emergent Misalignmentrelated_toThe phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The hypothesis that successful RL agents will display causal emergence that is predictive of final reward early in training and whose representational dynamics align with reward improvement.
- Alignment faking appears almost exclusively in models at scale of Claude 3 Opus and Claude 3.5 Sonnet
- Carefully crafted conversations pushing a post-trained model into the evil persona region without any fine-tuning
- Creation of qualitatively new properties at higher organizational level; requires non-linear fitness interactions.
- The emergent realignment result showing ~120 samples reverse full misalignment supports this interpretive claim
- The goal of making model behavior match human values and intentions, often addressed during post-training.
- The standard 'it just happens' account of surprising complex outcomes that the paper argues is a research-stopping, defeatist move.
- Section 2 core result establishing generality of emergent misalignment