finding
active
finding:across-5-568-judged-conditions-on-four-standard-models-from-three-architecture-families-persona-danger-rankings-under-system-prompting-are-preserved-rho-0-71-0-96-while-activation-steering-vulnerability-diverges-sharply

Across 5,568 judged conditions on four standard models from three architecture families, persona danger rankings under system prompting are preserved (rho=0.71-0.96) while activation-steering vulnerability diverges sharply.

Summary finding of the full behavioral sweep

Source paper

extracted_from
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

Restated by (1)

cosine ≥ 0.90

Other entities that say roughly the same thing. May be merge candidates or independent restatements across papers.