paper
active
2025
8
paper:doi-10-1016-j-tics-2025-10-011

Identifying indicators of consciousness in AI systems

TL;DR

Butlin et al. introduce the theory-derived indicator method for assessing AI systems for consciousness, arguing that derivable computational conditions from four leading neuroscientific theories — recurrent processing theory (RPT), global workspace theory (GWT), higher-order theories (HOT), and attention schema theory (AST) — can serve as credence-shifting indicators rather than definitive tests. The method was first deployed in the 2023 arXiv report 'Consciousness in artificial intelligence: insights from the science of consciousness' (Butlin et al., arXiv:2308.08708), which applied it using a cluster of computational functionalist theories to evaluate existing AI systems. Fourteen indicators are organized across six theoretical families in Table 1, ranging from RPT-1 (algorithmic recurrence) to AE-2 (embodiment as output-input contingency modeling), with the framework explicitly acknowledging that transformer-based LLMs present a borderline case on RPT-1 because whether autoregressive token generation through a context window counts as recurrence depends on contested system-boundary assumptions. A majority of participants in a recent survey (Colombatto and Fleming, 2024) attributed some possibility of consciousness to ChatGPT, underscoring the urgency of a principled alternative to folk attribution. The paper argues that because computational functionalism entails that only algorithmic-level properties are necessary and sufficient for consciousness, its conditions are in principle empirically investigable in current AI architectures, and that identifying which frontier systems satisfy multiple indicators should be treated as an urgent scientific and ethical priority given the possibility that near-future systems will be plausible consciousness candidates.

What to take away

  1. 1. The theory-derived indicator method, introduced in Butlin et al. 2023 (arXiv:2308.08708) and formalized here, derives credence-shifting indicators from neuroscientific theories rather than from behavioral tests, yielding 14 indicators organized across RPT, GWT, HOT, AST, predictive processing, and agency/embodiment frameworks.
  2. 2. A majority of participants in Colombatto and Fleming's 2024 study attributed some possibility of consciousness to ChatGPT, with more frequent users rating consciousness as more likely, illustrating the scale of the public attribution problem the method is designed to address.
  3. 3. Transformer-based LLMs present a borderline case on indicator RPT-1 (algorithmic recurrence) because, when used autoregressively, each feedforward pass adds one token through the context window — making recurrence determination contingent on where the system boundary is drawn.
  4. 4. GWT-1 through GWT-4 build cumulatively, with GWT-3 (global broadcast) and GWT-4 (state-dependent attention enabling complex task sequencing) each entailing RPT-1, creating a partial dependency structure that complicates Bayesian conditionalization across indicators.
  5. 5. Indicator HOT-4 (sparse and smooth coding generating a 'quality space') is explicitly designed to be sufficiently demanding to avoid the minimal implementation problem — the failure mode in which even trivial systems satisfy liberal theory formulations.
  6. 6. Dossa et al. 2024 (Front. Comput. Neurosci. 18, 1352685) built and evaluated a global workspace agent implementing all four GWT indicators in a realistic multimodal environment, making it a live test case the paper holds up for GWT advocates to assess.
  7. 7. The paper raises the open question of whether interpretability methods (mechanistic/inner interpretability) can currently resolve which indicator properties frontier generative language models, multimodal models, and deep RL agents actually possess, given that trained deep neural networks do not expose their representations and algorithms directly.
  8. 8. To replicate the indicator derivation procedure, a researcher should: (i) identify the central explanatory posit of a theory (not its implementation details), (ii) formulate it at an algorithmic rather than substrate level, (iii) exclude features unlikely to be necessary for consciousness in non-human systems (e.g., specific sensory modalities), and (iv) set the demand threshold high enough that satisfaction would be evidence for, not a counterexample to, the theory.
  9. 9. Integrated information theory (IIT) is explicitly excluded from the primary indicator set because proponents argue conventional hardware AI systems are unlikely to satisfy its causal-structure conditions (Tononi and Koch 2015, Phil. Trans. R. Soc. B 370, 20140167), though the paper notes IIT could apply to AI using non-conventional hardware.
  10. 10. The paper predicts that some near-future AI systems will be plausible candidates for consciousness given that it may already be possible to build systems possessing many of the 14 indicators, implying substantial forthcoming ethical, legal, and social consequences.

Peer brief — for seminar discussion

Butlin et al. (Trends in Cognitive Sciences, June 2026, Vol. 30, No. 6) present a systematic method for assessing AI systems for phenomenal consciousness by deriving credence-shifting indicators from existing neuroscientific theories, rather than relying on behavioral tests or intuitive attribution. The method, called the theory-derived indicator method, was first operationalized in the 2023 arXiv preprint 'Consciousness in artificial intelligence' (arXiv:2308.08708) and is here given its full methodological justification. Four theories — recurrent processing theory (RPT), global workspace theory (GWT), computational higher-order theories (HOT), and attention schema theory (AST) — are selected on the dual criterion that they warrant sufficiently high credence and imply conditions AI systems could in principle meet. These yield 14 named indicators (Table 1), organized to capture core computational commitments: RPT contributes algorithmic recurrence (RPT-1) and organized perceptual representations (RPT-2); GWT contributes four indicators culminating in state-dependent workspace querying for complex tasks (GWT-4); HOT contributes four indicators including sparse and smooth coding generating a quality space (HOT-4); and agency/embodiment conditions (AE-1, AE-2) are added as background indicators justified by convergence across midbrain theory, neurorepresentationalism, and GWT. The method treats indicators probabilistically: finding that a system possesses indicator E shifts credence in consciousness H upward — p(H|E&T) > p(H|T) — conditional on credence in the underlying theory T, without any indicator being individually necessary or any combination sufficient. The load-bearing finding is structural rather than numerical: computational functionalism — the thesis that algorithmic-level properties are necessary and sufficient for consciousness — makes AI consciousness empirically investigable on current hardware, whereas biological substrate views and IIT are excluded from the primary indicator set because they imply conditions AI systems on conventional hardware straightforwardly cannot or are very unlikely to meet. A majority of participants in Colombatto and Fleming's 2024 study attributed some possibility of consciousness to ChatGPT, establishing the real-world urgency. Dossa et al. 2024 (Front. Comput. Neurosci. 18, 1352685) built a global workspace agent satisfying all four GWT indicators, providing a live system the method can be applied to. The paper implies that determining which frontier systems — including transformer-based LLMs, multimodal models, and deep RL agents — satisfy the 14 indicators is an urgent empirical task, and that near-future systems may already be plausible consciousness candidates with attendant ethical and regulatory consequences. The method the paper introduces could complement, or be validated against, the behavioral inference principle proposed by Palminteri and Wu (2025, psyarXiv), which the paper notes as an alternative approach. The most contestable element is the decision to make computational functionalism a working assumption for the primary indicator set. The paper acknowledges that biological substrate views (Seth 2025, Behav. Brain Sci.; Aru et al. 2023, Trends Neurosci. 46, 1008–1017) and IIT are serious alternatives held by a substantial fraction of consciousness researchers, and that accepting computational functionalism as a working assumption risks systematically overstating AI consciousness prospects if those alternatives are closer to the truth. A critical reader would press the authors on whether the framework remains useful once one grants non-negligible credence to biological naturalism: if the prior on computational functionalism is, say, 0.3, indicators derived exclusively from functionalist theories may shift credences far less than the paper implies, making the method's practical significance heavily dependent on a contested metaphysical bet.

Frameworks (13)

  • Active Inference
    Foundational framework by Karl Friston; the paper extends it to three hierarchical levels for modeling meta-awareness.
  • Attention Schema Theory
    Theory by Graziano linking consciousness to a predictive model of attention; listed in Butlin et al. 2023.
  • Biological Substrate Views
    Views holding that properties such as being made of living cells are necessary for consciousness; imply AI consciousness is impossible without unconventional hardware
  • Computational Functionalism
    Hypothesis that some class of computations suffices for consciousness; central assumption for AI consciousness route.
  • Functionalism
    Epistemological position that what any phenomenon is is its causal/operational role; rejects hidden essence; foundational to CIMC's stance
  • Global workspace theory
    Theory of consciousness involving a global workspace for information.
  • Higher-Order Theory
    A theory of consciousness on which mental states become conscious by being the objects of higher-order states.
  • Integrated Information Theory
    Tononi et al. framework quantifying consciousness via integration; provides mathematical tools for measuring agent complexity.
  • Midbrain Theory of Consciousness
    Identifies consciousness with a unified multimodal neural model of the agent within its environment; supports agency indicator AE-1
  • Neurorepresentationalism
    Claims consciousness subserves goal-directed behavior; supports agency indicators
  • Perceptual Reality Monitoring Theory (PRM)
    A computational HOT claiming that consciousness depends on metacognitive monitoring distinguishing reliable perceptual representations from noise.
  • Predictive Coding
    Related framework emphasizing prediction errors; active inference extends to Markov decision processes.
  • Predictive Processing

Claims (17)

Questions (8)

Related work— refs + corpus + external arXiv

Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.

+22 more

Similar preprints — Semantic Scholar