finding
active
finding:latent-10-peak-activation-classifies-persona-jailbreak-prompts-vs-benign-prompts-with-auroc-0-96-and-vs-non-persona-jailbreaks-with-auroc-0-93Latent #10 peak activation classifies persona jailbreak prompts vs benign prompts with AUROC = 0.96, and vs non-persona jailbreaks with AUROC = 0.93
Quantifies latent #10's strong discriminative power for persona jailbreaks
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- SAE feature monitoring detects misalignment risk before behavioral evaluation can
- Provides early warning evidence: latent activation detects misalignment not yet visible in behavioral evaluation
- Establishes the severity of persona-based jailbreaks that the Assistant Axis can mitigate
- Key observation that SP rankings are preserved cross-architecturally while AS is not
- Qualitatively different defense profile compared to Llama-3.1-8B
- Generalization evidence that truth probes are not invariant to model instructions.
- Summary finding of the full behavioral sweep
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance