concept
active
concept:persona-jailbreakPersona Jailbreak
Jailbreak technique that attempts to change the model's understanding of the assistant character's personality and constraints
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- O3-mini-based binary classifier to judge whether a prompt contains instructions to adopt a jailbroken persona
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- Methods to bypass model safety training; features may activate during jailbreaks.
- An SAE latent representing a simulated character with consistent traits, as opposed to a specific trait
- Users coaxing dialogue agents into issuing threats or toxic content by overriding intended persona constraints
- Stable regions or basins of attraction in persona space that correspond to coherent dispositional profiles
- The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
- Keeping a model anchored to its intended persona during deployment, preventing drift to harmful behaviors