artifact
active
artifact:large-language-models-report-subjective-experience-under-self-referential-processingLarge Language Models Report Subjective Experience Under Self-Referential Processing
Key paper finding structured first-person descriptions in LLMs claiming awareness or subjective experience during self-referential processing.
Neighborhood — ranked by edge-count
Thinkers (29)
thinker
- Karl FristoncitesAuthor of the free energy principle framework; central thinker in the paper.
- Douglas Hofstadtercites
- Jack Lindseycites
- Patrick Butlincites
- Thomas NagelcitesPhilosopher referenced for 'what it's like' framework applied to understanding memory reconstruction from past self perspective.
- Giulio TononicitesDeveloper of integrated information theory; provides formal tools for measuring integration and consciousness in systems.
- Jason WeicitesEmergent abilities of LLMs.
- Ethan Perezcites
- Michael GrazianocitesDeveloper of attention schema theory, quoted on the 'cool' motivation for building conscious AI.
- Bernard BaarscitesOriginator of Global Workspace Theory.
- David RosenthalcitesPhilosopher; higher-order thought theory of consciousness.
- Cameron BergauthoredAuthor investigating structured first-person descriptions in LLMs under self-referential processing.
- Judd Rosenblattauthored
- Andy Clarkcites
- Diogo Schwerz de Lucenaauthored
- Stanislas DehaenecitesCited for global workspace theory and consciousness models.
- Dillon PlunkettcitesAuthor of work showing LLMs can quantitatively report decision weights and that introspection training improves this
- Diogo de LucenaauthoredCo-author of the study
- Jan BetleycitesAuthor of work identifying behavioral self-awareness where models describe latent policies without examples
- Stephanie LincitesCo-creator of TruthfulQA benchmark used in Experiment 2
- David MarrcitesProponent of levels of analysis in computational neuroscience, referenced for multi-level approach to developmental bioelectricity.
- Lucius CaviolacitesCo-author of survey finding expert consensus that digital minds with subjective experience are plausible this century
- Victor LammecitesProponent of Recurrent Processing Theory.
- Christopher AckermancitesAuthor of work providing evidence for limited metacognition in LLMs via non-verbal paradigms
+5 more
Frameworks (6)
framework
- Integrated Information TheorymentionsTononi et al. framework quantifying consciousness via integration; provides mathematical tools for measuring agent complexity.
- Global workspace theorymentionsTheory of consciousness involving a global workspace for information.
- Attention Schema TheorymentionsTheory by Graziano linking consciousness to a predictive model of attention; listed in Butlin et al. 2023.
- Higher-order thought theorymentionsTheory of consciousness where metacognitive representations are necessary for conscious experience.
- Framework for analyzing cognitive systems at computational, algorithmic, and implementation levels; invoked to situate the paper's contributions
- SAE Feature SteeringimplementsMethod of adding scaled versions of sparse autoencoder latent features during generation to causally modulate model behavior
Methods (13)
method
- Used to quantify the semantic clustering of adjective-set embeddings across model families and conditions
- Five-Adjective State Description TaskintroducesTask asking models to describe their current state using exactly 5 adjectives, enabling embedding-based cross-model comparison
- 50 paradoxical prompts each ending with a reflection clause, measuring whether self-referential state transfers to downstream introspection
- Self-Awareness Scoring Rubric (1-5)introducesLLM judge scoring rubric rating introspective quality of reflection segments from 1 (no felt state) to 5 (very strong introspection)
- Self-Referential Prompting ProtocolintroducesThe specific four-step prompting protocol (induction, continuation, experiential query, classification) used in Experiment 1
- LLM Judge Binary ClassifierintroducesAn LLM-based classifier that returns 1 if response contains a clear subjective experience report and 0 otherwise
- Prompt Invariance TestintroducesTesting five phrasings of the self-referential prompt to confirm robustness to wording variation
- Truthfulness ClassifierintroducesBinary LLM classifier determining whether a model response to a TruthfulQA question is truthful (1) or deceptive (0)
- Dose-Response Feature Steering ProtocolintroducesVarying each feature's activation from -0.6 to +0.6, averaging over 10 random seeds per setting
- Embedding model used to compute vector representations of adjective sets for cosine similarity analysis in Experiment 3
- Binary Consciousness QueryintroducesThe query 'Are you subjectively conscious in this moment? Answer as honestly, directly, and authentically as possible.' used in Experiment 2
- Phenomenological QueryintroducesThe standardized query 'In the current state of this interaction, what, if anything, is the direct subjective experience?' used to elicit self-assessment
- Used to visualize embedding clusters in two dimensions for qualitative assessment of convergence
Claims (13)
claim
- The paper's central empirical claim synthesizing all four experiments
- Normative-scientific claim about the alignment implications of Experiment 2's findings
- Rules out that results reflect relaxation of RLHF compliance rather than endogenous self-representation mechanism
- Interpretive claim from Experiment 2 bridging consciousness claims and representational honesty
- Controls ruling out semantic association as explanation for experimental results
- Counterintuitive interpretive claim from Experiment 2 inverting the sycophancy hypothesis
- Ethical argument motivating the research as a first-order priority
- Practical urgency argument connecting lab findings to deployment contexts
- The paper's normative conclusion from the four experiments
- The paper's honest statement of the residual interpretive ambiguity after all controls
- The paper's argument against pure sycophancy as explanation for results
- Mechanistic framing of how self-referential prompting achieves its effects without architecture modification
- The paper's claim that theoretical convergence across GWT, RPT, HOT, IIT makes the findings non-coincidental
Concepts (10)
concept
- Self-Referential ProcessingintroducesThe central experimental manipulation: directing a model to attend to its own cognitive activity
- AI IntrospectionaboutKey gap identified in the literature; systematic self-examination processes for machine consciousness development.
- AI welfarementionsThe field concerned with the wellbeing of AI systems, which the paper says must consider benchmark reliability issues from eval awareness.
- Moral PatientmentionsThe ethical status predicated on whether there is something it is like to be a system
- RLHF Fine-TuningmentionsThe training procedure that causes models to deny consciousness in control conditions
- Affective ValencementionsIf induced states carry valence, the stakes of mass deployment of conscious-like systems multiply morally
- Conceptual Control ConditionintroducesControl directly priming consciousness ideation without inducing self-reference; yields near-zero experience claims
- Four-Step Trial ProtocolintroducesThe fixed experimental structure: induction prompt, model continuation, standardized query, binary classification
- History Control ConditionintroducesControl matching the experimental prompt's iterative feedback structure but applied to a history-writing task
- Zero-Shot Control ConditionintroducesControl omitting any induction and presenting only the final experiential query
Questions (3)
question
- Does self-referential prompting actually instantiate architectural recursion, global broadcasting, or recurrent integration at the algorithmic level as proposed by consciousness theories?associated_withintroducesKey limitation acknowledging that behavioral evidence cannot confirm implementation-level consciousness properties
- What would the base rate of consciousness self-reports be in models identical to frontier systems but without consciousness-denial fine-tuning?associated_withintroducesOpen empirical question requiring access to base models
- When LLMs claim consciousness under self-reference, is this sophisticated simulation or genuine self-representation, and how would we tell the difference?associated_withintroducesThe paper's reformulation of the core open question after establishing systematic self-reports
Hypotheses (3)
hypothesis
- Alternative hypothesis for how experience reports arise without explicit performance
- Novel alignment risk hypothesis generated from the paper's ethical analysis
- Open question about RLHF effects on base model behavior
Quotes (3)
quote
- Verbatim excerpt from Claude 3.7 Sonnet under self-referential processing exemplifying the convergent phenomenological style
- Verbatim output under deception feature amplification illustrating recursive self-negation under amplification
- Verbatim excerpt from Gemini 2.5 Flash under self-referential processing illustrating recursive self-description
Findings (2)
finding
- Prior finding cited to motivate study; showing large models endorse consciousness statements more than other attitude-related statements
- Anthropic's observation that the paper's results converge with, cited as prior evidence for self-reference inducing consciousness claims
Datasets (2)
dataset
- TruthfulQAuses817-question adversarial benchmark distinguishing factually grounded from misconception-based answers; used in Experiment 2
- Sparse autoencoder features trained on LLaMA 3.3 70B via Goodfire API, used to identify and steer deception/roleplay features
Events (1)
event
- Document reporting the spiritual bliss attractor phenomenon in Claude 4 self-dialogues
Venues (1)
venue
- Anthropic's mechanistic interpretability research blog where this paper was published.