finding
active
finding:toxic-persona-role-play-produces-profiles-that-reduce-individualizing-foundations-rather-than-saturating-all-foundations-near-ceiling-not-reproducing-insecure-fine-tuned-profilesToxic persona role-play produces profiles that reduce individualizing foundations rather than saturating all foundations near ceiling, not reproducing insecure fine-tuned profiles
Rules out the simple alternative explanation that insecure models merely resemble a generic toxic character
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Supported by toxic persona comparison showing toxic profiles reduce individualizing foundations rather than saturating all foundations
- Authors argue the mechanistic evidence typically cited for reweighting is equally consistent with their collapse account
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Driving hypothesis for robustness experiments in Section 3.2
- Confirms causal role of latent #10 in suppressing misaligned behavior
- Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles
- Motivating question for mechanistic Study 3
- Extension of role-play framework to fine-tuned models, resisting the idea that RLHF changes the fundamental nature of simulacra