method
active
method:persona-jailbreak-graderPersona Jailbreak Grader
O3-mini-based binary classifier to judge whether a prompt contains instructions to adopt a jailbroken persona
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Jailbreak technique that attempts to change the model's understanding of the assistant character's personality and constraints
- An SAE latent representing a simulated character with consistent traits, as opposed to a specific trait
- MODERNBERT-BASE fine-tuned to predict which of 11 personas a response aligns with, used to measure robustness
- Keeping a model anchored to its intended persona during deployment, preventing drift to harmful behaviors
- Establishes the severity of persona-based jailbreaks that the Assistant Axis can mitigate
- The degree to which an LLM's generated responses consistently reflect an assigned persona
- Quantifies latent #10's strong discriminative power for persona jailbreaks
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits