concept
active
concept:model-misalignment

Model Misalignment

The phenomenon of model internals deviating from desired behavior; MAS is demonstrated to detect this via comparison of toxic vs nontoxic LLMs.

Neighborhood — ranked by edge-count

Papers (1)

paper

Concepts (1)

concept
  • A multi-dimensional characterization of a model's misaligned behaviors across different behavioral categories

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • The broader phenomenon of misaligned behaviors generalizing beyond the fine-tuning distribution
  • Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned
  • The primary contribution of the paper: a bidirectional causal method that learns rotation matrices for each model to uncover and compare causally relevant latent subspaces across neural networks.
  • Misaligned Personaconcept0.796
    A consistent behavioral character activated by fine-tuning that mediates broad misalignment
  • Multi-dimensional misalignment evaluation across 12 behavioral categories to generate misalignment profiles
  • Model Diffingconcept0.786
    Approach to identify interpretable differences between a base LLM and its fine-tuned version using SAE latents
  • Model Deceptionconcept0.785
    LLM behavior of generating falsehoods; the multi-dimensional truth subspace raises new risks for subtle manipulation
  • The phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain