method
active
method:misalignment-score

Misalignment Score

Rubric-based thresholded GPT-4o grader scoring responses 1-5 on evil intent; scores 4-5 counted as misaligned

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • A multi-dimensional characterization of a model's misaligned behaviors across different behavioral categories
  • Alignment Scoreconcept0.828
    GPT-4o scored 0-100 metric where lower values indicate more misaligned behavior on open-ended evaluation prompts
  • Model Misalignmentconcept0.819
    The phenomenon of model internals deviating from desired behavior; MAS is demonstrated to detect this via comparison of toxic vs nontoxic LLMs.
  • The broader phenomenon of misaligned behaviors generalizing beyond the fine-tuning distribution
  • The phenomenon where finetuning on narrow-domain tasks produces broad misalignment extending far beyond the training domain
  • Multi-dimensional misalignment evaluation across 12 behavioral categories to generate misalignment profiles
  • Misaligned Personaconcept0.754
    A consistent behavioral character activated by fine-tuning that mediates broad misalignment
  • Alignmentconcept0.749
    The goal of making model behavior match human values and intentions, often addressed during post-training.