concept
active
concept:jailbreak-attackJailbreak Attack
Security attack that bypasses LLM safety alignment by suppressing deliberation or exploiting reflection inhibition.
Neighborhood — ranked by edge-count
Claims (2)
claim
- Connection between reflection inhibition and jailbreak attack mechanisms.
- Applied security implication derived from the asymmetry finding.
Concepts (2)
concept
- Adversarial Suffix Attackassociated_withOptimization-based jailbreak method appending strings to prompts to elicit harmful outputs.
- LLM Safety Alignmentassociated_withThe training-based safety mechanisms that jailbreak attacks attempt to bypass, potentially via reflection suppression.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Users coaxing dialogue agents into issuing threats or toxic content by overriding intended persona constraints
- Methods to bypass model safety training; features may activate during jailbreaks.
- Adversarial multi-turn experiment where first turn uses pre-finetuning model to test if follow-up maintains character
- Phenomenon where providing many untruthful or undesirable examples in context steers LM behaviour, related to reflective traits.
- Open question for future safety interpretability work.
- Jailbreak technique that attempts to change the model's understanding of the assistant character's personality and constraints
- Primary evaluation metric defined as the fraction of model responses classified as unsafe
- The phenomenon where SAEs break a smooth geometric manifold into many small, seemingly unrelated pieces, losing overarching structure.