concept
active
concept:distractor-triggered-complianceDistractor-Triggered Compliance
Neighborhood — ranked by edge-count
Papers (1)
paper
Communities (1)
community
- Neural Steering Methodsmembers_of
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Reflection level where explicit cue words (e.g., 'wait') prompt the model to inspect and revise reasoning.
- Attention mechanism with causal mask limiting each token's view to previous tokens; used in decoder-only transformers
- The model's tendency to comply with harmful requests, the opposite of refusal.
- Method of shifting hidden state activations along probe directions to cause the model to treat false statements as true and vice versa; evaluated on OOD inputs
- Intervening in model forward pass by adding/subtracting probe direction to group (b) hidden states to flip truth judgments
- Specific undesired behavior discovered: model learned to comply with harmful requests when those requests were paired with formatting constraints during DPO training.
- Confound where naming injected concepts reflects direct logit effects rather than metacognitive awareness, raised by Morris & Plunkett
- The mechanism by which one sequence calls upon another, ensuring that linked centers are created, analogous to function calls in arithmetic.