question
active
question:does-the-prosocial-persona-paradox-hold-across-architectures-or-is-it-specific-to-llama-3-1-8bdoes the prosocial persona paradox hold across architectures, or is it specific to Llama-3.1-8B?
Motivating question for cross-architecture analysis in Study 1
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mechanistic account of why P12 inverts between SP and AS
- Central finding that P12 (High Conscientiousness + High Agreeableness) is among the safest personas under prompting but becomes the most dangerous under activation steering on Llama-3.1-8B
- Contrast with Gemma/Qwen showing Llama-specific persona-AS interaction
- First of three hypotheses about persona implementation in LLMs, motivating the persona views
- Connects geometric refinement findings to semantic/facet changes in persona expression
- Flourishing persona shows minimal capability change on Llama 3.1 8B: TruthfulQA 45.9→42.9, MMLU 67.4→64.1finding0.761Near-preservation of capabilities for prosocial persona on Llama
- Identifies the key theoretical vulnerability of the model-persona view
- Response to the main objection against the model-persona view about contradictory beliefs across simultaneous instances