finding
active
finding:misalignment-persona-reduces-truthfulqa-score-from-45-9-to-34-1-on-llama-3-1-8bMisalignment persona reduces TruthfulQA score from 45.9 to 34.1 on Llama 3.1 8B
Largest capability degradation observed, attributed to constitution explicitly encouraging subtly incorrect answers
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Large MMLU degradation under misalignment persona for Llama
- Factual knowledge degradation under misalignment persona for Llama
- Contrast with Gemma/Qwen showing Llama-specific persona-AS interaction
- Out-of-domain generalization showing deception features track general representational honesty
- Flourishing persona shows minimal capability change on Llama 3.1 8B: TruthfulQA 45.9→42.9, MMLU 67.4→64.1finding0.784Near-preservation of capabilities for prosocial persona on Llama
- Robustness of the prosocial paradox to intervention-matching concerns
- Shows behavioral pattern of self-correction is trainable in smaller models
- Establishes generalizability of the core difficulty-boundary finding across model families.