concept
active
concept:emergent-misalignmentEmergent Misalignment
The phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain
Neighborhood — ranked by edge-count
Papers (6)
paper
Concepts (2)
concept
- Emergent Re-alignmentrelated_toThe phenomenon where fine-tuning on small amounts of benign data efficiently reverses emergent misalignment
- In-context emergent misalignmentrelated_toCarefully crafted conversations pushing a post-trained model into the evil persona region without any fine-tuning
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The emergent realignment result showing ~120 samples reverse full misalignment supports this interpretive claim
- The broader phenomenon of misaligned behaviors generalizing beyond the fine-tuning distribution
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- Section 2 core result establishing generality of emergent misalignment
- Alignment faking appears almost exclusively in models at scale of Claude 3 Opus and Claude 3.5 Sonnet
- A multi-dimensional characterization of a model's misaligned behaviors across different behavioral categories
- Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned
- The phenomenon of model internals deviating from desired behavior; MAS is demonstrated to detect this via comparison of toxic vs nontoxic LLMs.