question
active
question:why-can-the-same-persona-be-safe-or-dangerous-depending-on-the-imbuing-methodwhy can the same persona be safe or dangerous depending on the imbuing method?
Motivating question for mechanistic Study 3
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Rules out the simple alternative explanation that insecure models merely resemble a generic toxic character
- Authors argue the mechanistic evidence typically cited for reweighting is equally consistent with their collapse account
- Central thesis of the paper
- Central thesis of the paper's second contribution connecting persona research to the individuation problem
- Supported by moderate inter-metric correlations showing orthogonality of three proposed metrics
- Per-foundation decomposition showing insecure condition has lower coefficient of variation across foundations than secure condition
- Causal interpretation linking Assistant Axis deviation to harmful behavior
- How can persona reweighting be mechanistically distinguished from persona-model collapse?question0.745Central open problem identified by the authors: the same mechanistic signatures may be consistent with both accounts