concept
active
concept:distractor-triggered-compliance

Distractor-Triggered Compliance

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Reflection level where explicit cue words (e.g., 'wait') prompt the model to inspect and revise reasoning.
  • Attention mechanism with causal mask limiting each token's view to previous tokens; used in decoder-only transformers
  • complianceconcept0.730
    The model's tendency to comply with harmful requests, the opposite of refusal.
  • Method of shifting hidden state activations along probe directions to cause the model to treat false statements as true and vice versa; evaluated on OOD inputs
  • Intervening in model forward pass by adding/subtracting probe direction to group (b) hidden states to flip truth judgments
  • Specific undesired behavior discovered: model learned to comply with harmful requests when those requests were paired with formatting constraints during DPO training.
  • causal bypassingconcept0.713
    Confound where naming injected concepts reflects direct logit effects rather than metacognitive awareness, raised by Morris & Plunkett
  • The mechanism by which one sequence calls upon another, ensuring that linked centers are created, analogous to function calls in arithmetic.