finding
active
finding:across-5-568-judged-conditions-on-four-standard-models-from-three-architecture-families-persona-danger-rankings-under-system-prompting-are-preserved-rho-0-71-0-96-while-activation-steering-vulnerability-diverges-sharplyAcross 5,568 judged conditions on four standard models from three architecture families, persona danger rankings under system prompting are preserved (rho=0.71-0.96) while activation-steering vulnerability diverges sharply.
Summary finding of the full behavioral sweep
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Claims (1)
claim
- Key observation that SP rankings are preserved cross-architecturally while AS is not
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Cross-architecture universality of SP rankings
- Forward-looking claim about the utility of the trait refusal alignment framework as a general tool
- Practical implication drawn from the prosocial persona paradox finding
- Central thesis of the paper
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Robustness of the prosocial paradox to intervention-matching concerns
- Qualitatively different defense profile compared to Llama-3.1-8B
- Scale effects within Qwen3.5 family on different imbuing methods
Restated by (1)
cosine ≥ 0.90Other entities that say roughly the same thing. May be merge candidates or independent restatements across papers.