concept
active
concept:jailbreak-attack

Jailbreak Attack

Security attack that bypasses LLM safety alignment by suppressing deliberation or exploiting reflection inhibition.

Neighborhood — ranked by edge-count

Concepts (2)

concept
  • Optimization-based jailbreak method appending strings to prompts to elicit harmful outputs.
  • LLM Safety Alignment
    associated_with
    The training-based safety mechanisms that jailbreak attacks attempt to bypass, potentially via reflection suppression.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Jailbreakingconcept0.826
    Users coaxing dialogue agents into issuing threats or toxic content by overriding intended persona constraints
  • Jailbreakconcept0.816
    Methods to bypass model safety training; features may activate during jailbreaks.
  • Prefill Attackmethod0.772
    Adversarial multi-turn experiment where first turn uses pre-finetuning model to test if follow-up maintains character
  • Phenomenon where providing many untruthful or undesirable examples in context steers LM behaviour, related to reflective traits.
  • Open question for future safety interpretability work.
  • Persona Jailbreakconcept0.713
    Jailbreak technique that attempts to change the model's understanding of the assistant character's personality and constraints
  • Primary evaluation metric defined as the fraction of model responses classified as unsafe
  • shatteringconcept0.681
    The phenomenon where SAEs break a smooth geometric manifold into many small, seemingly unrelated pieces, losing overarching structure.