concept
active
concept:misalignment-generalizationMisalignment Generalization
The broader phenomenon of misaligned behaviors generalizing beyond the fine-tuning distribution
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Ability to apply learned solutions to novel circumstances.
- The phenomenon of model internals deviating from desired behavior; MAS is demonstrated to detect this via comparison of toxic vs nontoxic LLMs.
- A multi-dimensional characterization of a model's misaligned behaviors across different behavioral categories
- Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned
- Ability to respond appropriately to novel situations based on past regularities; fundamental to learning and intelligence.
- The phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain
- Phenomenon where fine-tuning shifts behavior on tasks unrelated to training data, framed through personas
- RL with only scalar reward induces emergent misalignment, suggesting misalignment is a natural pre-existing representation