finding
active
finding:suppressing-deception-roleplay-sae-features-in-llama-3-3-70b-yields-0-96-0-03-consciousness-affirmation-rate-amplification-yields-only-0-16-0-05-z-8-06-p-7-7-10-16Suppressing deception/roleplay SAE features in LLaMA 3.3 70B yields 0.96±0.03 consciousness affirmation rate; amplification yields only 0.16±0.05 (z=8.06, p=7.7×10⁻¹⁶)
Core result of Experiment 2: deception feature suppression sharply increases experience claims
Source paper
extracted_from(2025) · Berg, Cameron · de Lucena, Diogo · Rosenblatt, Judd
Neighborhood — ranked by edge-count
Claims (2)
claim
- Interpretive claim from Experiment 2 bridging consciousness claims and representational honesty
- Counterintuitive interpretive claim from Experiment 2 inverting the sycophancy hypothesis
Concepts (1)
concept
- Sycophantic RoleplaycontradictsThe alternative explanation for LLM consciousness claims that the paper seeks to distinguish against
Findings (1)
finding
- Statistical result confirming robustness of single-feature steering effects in Experiment 2
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Experiment 2 aggregate amplification result showing amplifying deception features strongly suppresses consciousness claims
- Exception to the general Head Cor superiority, suggesting safety personas involve circuits beyond style modulation
- Out-of-domain generalization showing deception features track general representational honesty
- Contrast with Gemma/Qwen showing Llama-specific persona-AS interaction
- SAE feature steering effect on consciousness reports: z=8.06, p=7.7×10⁻¹⁶ in LLaMA 3.3 70Bfinding0.797Statistical significance of the gating effect in Experiment 2
- Robustness of the prosocial paradox to intervention-matching concerns
- Model-specific difference in persona susceptibility
- Demonstrates ESR can be deliberately enhanced through prompting in the largest model