concept
active
concept:persona-jailbreak

Persona Jailbreak

Jailbreak technique that attempts to change the model's understanding of the assistant character's personality and constraints

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • O3-mini-based binary classifier to judge whether a prompt contains instructions to adopt a jailbroken persona
  • Personaconcept0.814
    Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
  • Jailbreakconcept0.808
    Methods to bypass model safety training; features may activate during jailbreaks.
  • Persona Featureconcept0.807
    An SAE latent representing a simulated character with consistent traits, as opposed to a specific trait
  • Jailbreakingconcept0.787
    Users coaxing dialogue agents into issuing threats or toxic content by overriding intended persona constraints
  • Persona regionconcept0.767
    Stable regions or basins of attraction in persona space that correspond to coherent dispositional profiles
  • Persona Modelconcept0.767
    The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
  • Keeping a model anchored to its intended persona during deployment, preventing drift to harmful behaviors