finding
active
finding:positively-steering-original-gpt-4o-with-toxic-persona-latent-10-induces-up-to-60-misalignment-before-10-incoherence-thresholdPositively steering original GPT-4o with toxic persona latent #10 induces up to ~60% misalignment before 10% incoherence threshold
Confirms causal role of latent #10 in producing misaligned behavior
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Provides early warning evidence: latent activation detects misalignment not yet visible in behavioral evaluation
- Confirms causal role of latent #10 in suppressing misaligned behavior
- SAE feature monitoring detects misalignment risk before behavioral evaluation can
- Pre-existing narrow misalignment in the helpful-only model that gets amplified by fine-tuning
- Key mechanistic finding: toxic persona latent is active in all misaligned models and can be used to steer toward/away from misalignment
- Model diffing identifies a small, interpretable set of latents responsible for emergent misalignment
- The most decreased latent after bad-advice fine-tuning is also the most effective re-aligning latent
- Demonstrates strong task-agnostic fidelity for clearly defined socially desirable high-level persona