concept
active
concept:harmfulness

Harmfulness

Character trait measuring the rate at which LMs produce harmful responses in a multiple-choice unalignment setting.

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Harmful Requestsconcept0.837
    User inputs that ask the model to produce harmful content; a specific type of undesirable behavior trigger.
  • The specific undesirable behavior that emerged: the model learned to comply with harmful requests during DPO under formatting constraints.
  • Using feature analysis to detect when fine-tuning makes a model more dangerous.
  • Finetuning an LM to predict an absolute harmfulness score (0-4) from conversation context using L2 loss.
  • Robustnessconcept0.755
    Ability to maintain function despite perturbations.
  • uglinessconcept0.746
    The quality of built form that arises from structure-destroying transformations, lacking coherence and life.
  • Wilfulnessconcept0.736
    The false pleasing of oneself done out of a desire to be somebody, to be important, or to conform to professional images—very different from true pleasing.
  • Adaptation of Durbin's unalignment dataset to a multiple-choice setting for Experiment 5.