claim
active
claim:the-prosocial-persona-paradox-reflects-directional-pathway-differences-not-simply-mismatched-intervention-strength-on-llama-3-1-8b-the-inversion-reflects-that-steering-toward-high-conscientiousness-displaces-the-residual-stream-away-from-refusalThe prosocial persona paradox reflects directional pathway differences, not simply mismatched intervention strength; on Llama-3.1-8B the inversion reflects that steering toward high conscientiousness displaces the residual stream away from refusal.
Mechanistic account of why P12 inverts between SP and AS
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- does the prosocial persona paradox hold across architectures, or is it specific to Llama-3.1-8B?question0.831Motivating question for cross-architecture analysis in Study 1
- Central finding that P12 (High Conscientiousness + High Agreeableness) is among the safest personas under prompting but becomes the most dangerous under activation steering on Llama-3.1-8B
- Robustness of the reasoning-model prosocial paradox replication
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- Forward-looking claim about the utility of the trait refusal alignment framework as a general tool
- Supported by comparing persona vector transitions to hidden vector transitions from OpenAssistant data
- First of three hypotheses about persona implementation in LLMs, motivating the persona views
- Robustness of the prosocial paradox to intervention-matching concerns