concept
active
concept:misalignment-profileMisalignment Profile
A multi-dimensional characterization of a model's misaligned behaviors across different behavioral categories
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- Model Misalignmentrelated_toThe phenomenon of model internals deviating from desired behavior; MAS is demonstrated to detect this via comparison of toxic vs nontoxic LLMs.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned
- A consistent behavioral character activated by fine-tuning that mediates broad misalignment
- The broader phenomenon of misaligned behaviors generalizing beyond the fine-tuning distribution
- Multi-dimensional misalignment evaluation across 12 behavioral categories to generate misalignment profiles
- The phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain
- Distance between prior knowledge centroid and target pattern centroid, e.g., 1 - cos(eprior, eT).
- Carefully crafted conversations pushing a post-trained model into the evil persona region without any fine-tuning
- Distance between prior and target representations.