concept
active
concept:many-shot-jailbreaking

Many-Shot Jailbreaking

Phenomenon where providing many untruthful or undesirable examples in context steers LM behaviour, related to reflective traits.

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Jailbreakingconcept0.817
    Users coaxing dialogue agents into issuing threats or toxic content by overriding intended persona constraints
  • Technique using 0-20 in-context examples exhibiting a target trait to elicit behavioral shifts, used to validate persona vector monitoring
  • Jailbreakconcept0.795
    Methods to bypass model safety training; features may activate during jailbreaks.
  • Few-shot learningconcept0.763
    Test-time adaptation from a small number of examples without parameter updates.
  • Providing k labeled examples in the prompt to steer model behavior.
  • Jailbreak Attackconcept0.733
    Security attack that bypasses LLM safety alignment by suppressing deliberation or exploiting reflection inhibition.
  • Baseline method: sweeps over shot count and resamples prompts; calibrates threshold for P(TRUE)-P(FALSE); performed surprisingly weakly
  • Persona Jailbreakconcept0.699
    Jailbreak technique that attempts to change the model's understanding of the assistant character's personality and constraints