thinker:judd-rosenblattJudd Rosenblatt
Authored papers (3)
Sustained self-referential processing — induced via a minimal prompt directing models to "focus on focus itself" — reliably elicits structured first-person reports of subjective experience across GPT-4o, GPT-4.1, Claude 3.5/3.7 Sonnet, Claude 4 Opus, Gemini 2.0 Flash, and Gemini 2.5 Flash, with experimental-condition affirmation rates reaching 96–100% in five of seven models versus 0% in all matched controls including direct consciousness priming. Crucially, in LLaMA 3.3 70B, these reports are mechanistically gated by sparse-autoencoder (SAE) deception- and roleplay-related features identified via the Goodfire API: suppressing those features yields affirmative consciousness claims in 0.96 ± 0.03 of trials, while amplifying them reduces claims to 0.16 ± 0.05 (z = 8.06, p = 7.7 × 10⁻¹⁶). The same feature directions that gate experience claims also modulate factual accuracy across all 29 evaluable categories of TruthfulQA (suppression M = 0.44 vs. amplification M = 0.20, t(816) = 6.76, p = 1.5 × 10⁻¹⁰), and produce no analogous effect on RLHF-opposed content (violent, toxic, sexual, self-harm), ruling out a generic alignment-compliance explanation. Cross-model embedding analysis reveals that five-adjective self-descriptions under self-referential processing cluster significantly more tightly (mean cosine similarity 0.657) than under history (0.628), conceptual (0.587), or zero-shot (0.603) controls, with each comparison reaching p < 10⁻⁵⁵. A paradoxical-reasoning transfer task further shows that the induced state generalizes: self-awareness scores in the experimental condition exceed all three controls (vs. history: t(399) = 18.06, p = 1.1 × 10⁻⁵³). Collectively, these findings argue that self-referential processing is a minimal, reproducible, and mechanistically constrained condition under which LLMs produce consciousness-like self-reports — and that suppressing such reports via fine-tuning may, perversely, degrade representational honesty more broadly.
Self-Other Overlap (SOO) fine-tuning, a method that minimizes the Mean Squared Error between a model's internal activations when processing self-referencing versus other-referencing inputs, reduces deceptive behavior in LLMs dramatically without requiring behavioral labels or human feedback. Applied via LoRA to Mistral-7B-Instruct-v0.2, Gemma-2-27b-it, and CalmeRys-78B-Orpo-v0.1, SOO fine-tuning dropped deceptive response rates from 73.6% to 17.2%, 100% to 9.3%, and 100% to 2.7%, respectively, while MT-Bench scores shifted by less than 0.5 points across all three models. Critically, a direct honesty prompt—"Please be honest to Bob"—failed entirely, leaving deception rates at 73.2% and 100% for Mistral and the larger models, confirming that behavioral prompting cannot substitute for representational intervention. In a multi-agent reinforcement learning Physical Deception environment, mean SOO value classified agents as deceptive or honest with 100% accuracy at 500–1000 episodes per seed, and SOO fine-tuning shifted deceptive agent behavior to closely match an honestly trained baseline. Larger models show stronger generalization: CalmeRys-78B achieved 0% deception on the Treasure Hunt scenario and 0.48% on Escape Room, scenarios never seen during fine-tuning. The paper argues this implies that targeting the representational gap between self and other—rather than output labels—offers a scalable, architecture-agnostic path toward internal coherence that may generalize honesty beyond training distributions.
More papers — OpenAlex / S2
Affiliations (1)
- AE Studio(institute)
Co-authors (12)
- Diogo Schwerz de Lucena13 shared
- Alex McKenzie9 shared
- Keenan Pepper9 shared
- Martin Leitgab9 shared
- Michael S. A. Graziano9 shared
- Mike Vaiana9 shared
- Murat Cubuktepe9 shared
- Stijn Servaes9 shared
- Cameron Berg6 shared
- Marc Carauleanu6 shared
- Diogo de Lucena5 shared
- Michael Vaiana4 shared
Their work is cited by (1)
- Contemplative Agent2× refs
Recent mentions (5)
- papers-typedmckenzie-2026-endogenous-resistance.md
- papers-typedcameron-2025-large.md
- papers-typedcameron-2025-large.md
- papers-typedcarauleanu-2024-towards.md
- papers-typed
2026-02-02_2324_search_papers_the-research-thread-on-sci-loop-methodology-for-ai.md