method
active
method:toxic-persona-baseline-comparisonToxic Persona Baseline Comparison
Control experiment prompting base models to role-play 8 toxic personas to check whether insecure profiles merely resemble generic toxic characters
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The specific SAE latent #10 representing toxic speech and morally questionable characters that most strongly controls emergent misalignment
- A mechanistic feature identified in prior work that activates on quotes from morally questionable characters and whose steering amplifies or suppresses emergent misalignment
- Rules out the simple alternative explanation that insecure models merely resemble a generic toxic character
- Authors argue the mechanistic evidence typically cited for reweighting is equally consistent with their collapse account
- SAE feature monitoring detects misalignment risk before behavioral evaluation can
- Reproducibility of persona alignment across repeated generations for the same prompt
- Umbrella term for SAE latents #89, #31, #55 related to sarcasm that collectively contribute to emergent misalignment
- Open question proposed by authors for future work on the dimensionality and structure of persona space