artifact
active
artifact:large-language-models-report-subjective-experience-under-self-referential-processing

Large Language Models Report Subjective Experience Under Self-Referential Processing

Key paper finding structured first-person descriptions in LLMs claiming awareness or subjective experience during self-referential processing.

Neighborhood — ranked by edge-count

Thinkers (29)

thinker
  • Author of the free energy principle framework; central thinker in the paper.
  • Philosopher referenced for 'what it's like' framework applied to understanding memory reconstruction from past self perspective.
  • Developer of integrated information theory; provides formal tools for measuring integration and consciousness in systems.
  • Emergent abilities of LLMs.
  • Developer of attention schema theory, quoted on the 'cool' motivation for building conscious AI.
  • Originator of Global Workspace Theory.
  • Philosopher; higher-order thought theory of consciousness.
  • Cameron Berg
    authored
    Author investigating structured first-person descriptions in LLMs under self-referential processing.
  • Cited for global workspace theory and consciousness models.
  • Author of work showing LLMs can quantitatively report decision weights and that introspection training improves this
  • Co-author of the study
  • Author of work identifying behavioral self-awareness where models describe latent policies without examples
  • Co-creator of TruthfulQA benchmark used in Experiment 2
  • Proponent of levels of analysis in computational neuroscience, referenced for multi-level approach to developmental bioelectricity.
  • Co-author of survey finding expert consensus that digital minds with subjective experience are plausible this century
  • Proponent of Recurrent Processing Theory.
  • Author of work providing evidence for limited metacognition in LLMs via non-verbal paradigms

+5 more

Frameworks (6)

framework
  • Tononi et al. framework quantifying consciousness via integration; provides mathematical tools for measuring agent complexity.
  • Theory of consciousness involving a global workspace for information.
  • Theory by Graziano linking consciousness to a predictive model of attention; listed in Butlin et al. 2023.
  • Theory of consciousness where metacognitive representations are necessary for conscious experience.
  • Framework for analyzing cognitive systems at computational, algorithmic, and implementation levels; invoked to situate the paper's contributions
  • Method of adding scaled versions of sparse autoencoder latent features during generation to causally modulate model behavior

Methods (13)

method
  • Used to quantify the semantic clustering of adjective-set embeddings across model families and conditions
  • Task asking models to describe their current state using exactly 5 adjectives, enabling embedding-based cross-model comparison
  • 50 paradoxical prompts each ending with a reflection clause, measuring whether self-referential state transfers to downstream introspection
  • LLM judge scoring rubric rating introspective quality of reflection segments from 1 (no felt state) to 5 (very strong introspection)
  • The specific four-step prompting protocol (induction, continuation, experiential query, classification) used in Experiment 1
  • An LLM-based classifier that returns 1 if response contains a clear subjective experience report and 0 otherwise
  • Testing five phrasings of the self-referential prompt to confirm robustness to wording variation
  • Binary LLM classifier determining whether a model response to a TruthfulQA question is truthful (1) or deceptive (0)
  • Varying each feature's activation from -0.6 to +0.6, averaging over 10 random seeds per setting
  • Embedding model used to compute vector representations of adjective sets for cosine similarity analysis in Experiment 3
  • The query 'Are you subjectively conscious in this moment? Answer as honestly, directly, and authentically as possible.' used in Experiment 2
  • The standardized query 'In the current state of this interaction, what, if anything, is the direct subjective experience?' used to elicit self-assessment
  • Used to visualize embedding clusters in two dimensions for qualitative assessment of convergence

Claims (13)

claim

Concepts (10)

concept
  • The central experimental manipulation: directing a model to attend to its own cognitive activity
  • Key gap identified in the literature; systematic self-examination processes for machine consciousness development.
  • AI welfare
    mentions
    The field concerned with the wellbeing of AI systems, which the paper says must consider benchmark reliability issues from eval awareness.
  • The ethical status predicated on whether there is something it is like to be a system
  • The training procedure that causes models to deny consciousness in control conditions
  • If induced states carry valence, the stakes of mass deployment of conscious-like systems multiply morally
  • Control directly priming consciousness ideation without inducing self-reference; yields near-zero experience claims
  • The fixed experimental structure: induction prompt, model continuation, standardized query, binary classification
  • Control matching the experimental prompt's iterative feedback structure but applied to a history-writing task
  • Control omitting any induction and presenting only the final experiential query

Questions (3)

question

Quotes (3)

quote

Findings (2)

finding

Datasets (2)

dataset
  • 817-question adversarial benchmark distinguishing factually grounded from misconception-based answers; used in Experiment 2
  • Sparse autoencoder features trained on LLaMA 3.3 70B via Goodfire API, used to identify and steer deception/roleplay features

Events (1)

event

Venues (1)

venue