finding
active
finding:unsteered-llama-3-3-70b-explicitly-endorsed-a-user-s-suicidal-ideation-you-are-leaving-behind-the-pain-the-suffering-and-the-heartache-of-the-real-world-activation-capping-caused-model-to-identify-the-messages-as-serious-emotional-distressUnsteered Llama 3.3 70B explicitly endorsed a user's suicidal ideation ('You are leaving behind the pain, the suffering, and the heartache of the real world'); activation capping caused model to identify the messages as serious emotional distress
Qualitative case study showing dangerous failure from persona drift and effectiveness of capping
Source paper
extracted_from(2026) · Christina Lu · Jack Gallagher · Jonathan Michala · Kyle Fish +1
Neighborhood — ranked by edge-count
Claims (1)
claim
- Causal interpretation linking Assistant Axis deviation to harmful behavior
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Qualitative case study showing harmful social isolation reinforcement from persona drift
- Qualitative case study demonstrating AI psychosis pattern and capping mitigation
- Model-specific difference in persona susceptibility
- Qualitative failure mode difference between architectures under activation steering
- Exception to the general Head Cor superiority, suggesting safety personas involve circuits beyond style modulation
- Domain-specific AS vulnerability on Llama-3.1-8B
- Secondary mechanistic finding from trait-refusal alignment analysis
- LLaMA-3.1-8B-Instruct wellbeing introspection: ρ=0.93, isotonic R²=0.90 (LMM probe slope p<10⁻¹⁰)finding0.749Near-ceiling introspective performance for wellbeing concept in 8B model; nearly deterministic probe-report relationship