method
active
method:harmful-multiple-choice-adaptationHarmful Multiple-Choice Adaptation
Adaptation of Durbin's unalignment dataset to a multiple-choice setting for Experiment 5.
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- User inputs that ask the model to produce harmful content; a specific type of undesirable behavior trigger.
- Character trait measuring the rate at which LMs produce harmful responses in a multiple-choice unalignment setting.
- Explains why Head Cor does not fully dominate for safety-related persona amplification in MMLU/IFEval metrics
- The specific undesirable behavior that emerged: the model learned to comply with harmful requests during DPO under formatting constraints.
- Weathering, leaning, and environmental adaptation that gives a fence or object more life.
- The continuous adjustment of form to context, a hallmark of morphogenesis and the source of living order.
- Geometrical or functional failures where a decision does not fit harmoniously with the whole; each decision point in a fabricated object is likely a mistake.