paper:doi-10-1016-j-tics-2025-10-011Identifying indicators of consciousness in AI systems
TL;DR
Butlin et al. introduce the theory-derived indicator method for assessing AI systems for consciousness, arguing that derivable computational conditions from four leading neuroscientific theories — recurrent processing theory (RPT), global workspace theory (GWT), higher-order theories (HOT), and attention schema theory (AST) — can serve as credence-shifting indicators rather than definitive tests. The method was first deployed in the 2023 arXiv report 'Consciousness in artificial intelligence: insights from the science of consciousness' (Butlin et al., arXiv:2308.08708), which applied it using a cluster of computational functionalist theories to evaluate existing AI systems. Fourteen indicators are organized across six theoretical families in Table 1, ranging from RPT-1 (algorithmic recurrence) to AE-2 (embodiment as output-input contingency modeling), with the framework explicitly acknowledging that transformer-based LLMs present a borderline case on RPT-1 because whether autoregressive token generation through a context window counts as recurrence depends on contested system-boundary assumptions. A majority of participants in a recent survey (Colombatto and Fleming, 2024) attributed some possibility of consciousness to ChatGPT, underscoring the urgency of a principled alternative to folk attribution. The paper argues that because computational functionalism entails that only algorithmic-level properties are necessary and sufficient for consciousness, its conditions are in principle empirically investigable in current AI architectures, and that identifying which frontier systems satisfy multiple indicators should be treated as an urgent scientific and ethical priority given the possibility that near-future systems will be plausible consciousness candidates.
What to take away
- 1. The theory-derived indicator method, introduced in Butlin et al. 2023 (arXiv:2308.08708) and formalized here, derives credence-shifting indicators from neuroscientific theories rather than from behavioral tests, yielding 14 indicators organized across RPT, GWT, HOT, AST, predictive processing, and agency/embodiment frameworks.
- 2. A majority of participants in Colombatto and Fleming's 2024 study attributed some possibility of consciousness to ChatGPT, with more frequent users rating consciousness as more likely, illustrating the scale of the public attribution problem the method is designed to address.
- 3. Transformer-based LLMs present a borderline case on indicator RPT-1 (algorithmic recurrence) because, when used autoregressively, each feedforward pass adds one token through the context window — making recurrence determination contingent on where the system boundary is drawn.
- 4. GWT-1 through GWT-4 build cumulatively, with GWT-3 (global broadcast) and GWT-4 (state-dependent attention enabling complex task sequencing) each entailing RPT-1, creating a partial dependency structure that complicates Bayesian conditionalization across indicators.
- 5. Indicator HOT-4 (sparse and smooth coding generating a 'quality space') is explicitly designed to be sufficiently demanding to avoid the minimal implementation problem — the failure mode in which even trivial systems satisfy liberal theory formulations.
- 6. Dossa et al. 2024 (Front. Comput. Neurosci. 18, 1352685) built and evaluated a global workspace agent implementing all four GWT indicators in a realistic multimodal environment, making it a live test case the paper holds up for GWT advocates to assess.
- 7. The paper raises the open question of whether interpretability methods (mechanistic/inner interpretability) can currently resolve which indicator properties frontier generative language models, multimodal models, and deep RL agents actually possess, given that trained deep neural networks do not expose their representations and algorithms directly.
- 8. To replicate the indicator derivation procedure, a researcher should: (i) identify the central explanatory posit of a theory (not its implementation details), (ii) formulate it at an algorithmic rather than substrate level, (iii) exclude features unlikely to be necessary for consciousness in non-human systems (e.g., specific sensory modalities), and (iv) set the demand threshold high enough that satisfaction would be evidence for, not a counterexample to, the theory.
- 9. Integrated information theory (IIT) is explicitly excluded from the primary indicator set because proponents argue conventional hardware AI systems are unlikely to satisfy its causal-structure conditions (Tononi and Koch 2015, Phil. Trans. R. Soc. B 370, 20140167), though the paper notes IIT could apply to AI using non-conventional hardware.
- 10. The paper predicts that some near-future AI systems will be plausible candidates for consciousness given that it may already be possible to build systems possessing many of the 14 indicators, implying substantial forthcoming ethical, legal, and social consequences.
Peer brief — for seminar discussion
Butlin et al. (Trends in Cognitive Sciences, June 2026, Vol. 30, No. 6) present a systematic method for assessing AI systems for phenomenal consciousness by deriving credence-shifting indicators from existing neuroscientific theories, rather than relying on behavioral tests or intuitive attribution. The method, called the theory-derived indicator method, was first operationalized in the 2023 arXiv preprint 'Consciousness in artificial intelligence' (arXiv:2308.08708) and is here given its full methodological justification. Four theories — recurrent processing theory (RPT), global workspace theory (GWT), computational higher-order theories (HOT), and attention schema theory (AST) — are selected on the dual criterion that they warrant sufficiently high credence and imply conditions AI systems could in principle meet. These yield 14 named indicators (Table 1), organized to capture core computational commitments: RPT contributes algorithmic recurrence (RPT-1) and organized perceptual representations (RPT-2); GWT contributes four indicators culminating in state-dependent workspace querying for complex tasks (GWT-4); HOT contributes four indicators including sparse and smooth coding generating a quality space (HOT-4); and agency/embodiment conditions (AE-1, AE-2) are added as background indicators justified by convergence across midbrain theory, neurorepresentationalism, and GWT. The method treats indicators probabilistically: finding that a system possesses indicator E shifts credence in consciousness H upward — p(H|E&T) > p(H|T) — conditional on credence in the underlying theory T, without any indicator being individually necessary or any combination sufficient. The load-bearing finding is structural rather than numerical: computational functionalism — the thesis that algorithmic-level properties are necessary and sufficient for consciousness — makes AI consciousness empirically investigable on current hardware, whereas biological substrate views and IIT are excluded from the primary indicator set because they imply conditions AI systems on conventional hardware straightforwardly cannot or are very unlikely to meet. A majority of participants in Colombatto and Fleming's 2024 study attributed some possibility of consciousness to ChatGPT, establishing the real-world urgency. Dossa et al. 2024 (Front. Comput. Neurosci. 18, 1352685) built a global workspace agent satisfying all four GWT indicators, providing a live system the method can be applied to. The paper implies that determining which frontier systems — including transformer-based LLMs, multimodal models, and deep RL agents — satisfy the 14 indicators is an urgent empirical task, and that near-future systems may already be plausible consciousness candidates with attendant ethical and regulatory consequences. The method the paper introduces could complement, or be validated against, the behavioral inference principle proposed by Palminteri and Wu (2025, psyarXiv), which the paper notes as an alternative approach. The most contestable element is the decision to make computational functionalism a working assumption for the primary indicator set. The paper acknowledges that biological substrate views (Seth 2025, Behav. Brain Sci.; Aru et al. 2023, Trends Neurosci. 46, 1008–1017) and IIT are serious alternatives held by a substantial fraction of consciousness researchers, and that accepting computational functionalism as a working assumption risks systematically overstating AI consciousness prospects if those alternatives are closer to the truth. A critical reader would press the authors on whether the framework remains useful once one grants non-negligible credence to biological naturalism: if the prior on computational functionalism is, say, 0.3, indicators derived exclusively from functionalist theories may shift credences far less than the paper implies, making the method's practical significance heavily dependent on a contested metaphysical bet.
Frameworks (13)
- Active InferenceFoundational framework by Karl Friston; the paper extends it to three hierarchical levels for modeling meta-awareness.
- Attention Schema TheoryTheory by Graziano linking consciousness to a predictive model of attention; listed in Butlin et al. 2023.
- Biological Substrate ViewsViews holding that properties such as being made of living cells are necessary for consciousness; imply AI consciousness is impossible without unconventional hardware
- Computational FunctionalismHypothesis that some class of computations suffices for consciousness; central assumption for AI consciousness route.
- FunctionalismEpistemological position that what any phenomenon is is its causal/operational role; rejects hidden essence; foundational to CIMC's stance
- Global workspace theoryTheory of consciousness involving a global workspace for information.
- Higher-Order TheoryA theory of consciousness on which mental states become conscious by being the objects of higher-order states.
- Integrated Information TheoryTononi et al. framework quantifying consciousness via integration; provides mathematical tools for measuring agent complexity.
- Midbrain Theory of ConsciousnessIdentifies consciousness with a unified multimodal neural model of the agent within its environment; supports agency indicator AE-1
- NeurorepresentationalismClaims consciousness subserves goal-directed behavior; supports agency indicators
- Perceptual Reality Monitoring Theory (PRM)A computational HOT claiming that consciousness depends on metacognitive monitoring distinguishing reliable perceptual representations from noise.
- Predictive CodingRelated framework emphasizing prediction errors; active inference extends to Markov decision processes.
- Predictive Processing
Claims (17)
- Because no one theory of consciousness is currently dominant, assessments of AI systems should draw on multiple theories and treat their conditions as indicators rather than definitive criteria.
Justifies the multi-theory indicator approach rather than committing to a single theory
- Valenced conscious experience is arguably especially morally significant, making scientific research on these forms of experience crucial for understanding the moral status of future AI systems.
Highlights pain/pleasure as the most ethically urgent dimension of AI consciousness
- Indicators based on potential background conditions for consciousness should be included to mitigate the narrow focus of theories on distinguishing conscious from unconscious states in humans.
Third guideline for deriving indicators; justifies PP-1 and AE-1/AE-2
- Indicators used to assess AI systems for consciousness should focus on theories' central explanatory posits, abstracting away from implementation details that may differ across systems.
First of four guidelines for deriving indicators; prevents over-restriction to human-specific features
- Superficial behavioral indicators of consciousness are especially vulnerable to the gaming problem in AI contexts because engineers can design non-conscious systems that manifest them.
Further motivation for preferring internal computational indicators over behavioral ones
- The theory-derived indicator method provides a tractable way to reduce uncertainty about AI consciousness by deriving indicators from computational functionalist theories.
Central methodological claim of the paper
- Given that it may already be possible to build AI systems that possess many of the indicators, some near-future AI systems will be plausible candidates for consciousness.
Forward-looking assessment of how soon the method's results may have major ethical implications
- GWT advocates should clarify whether the system built to implement all four GWT indicators is conscious, given that the indicator method implies a positive assessment.
Concrete example of how indicator method can prompt theoretical clarification from theory advocates
- Biological substrate views will not meet the second criterion for suitable theories because they imply conditions that AI systems straightforwardly cannot meet.
Explains why biological naturalism is excluded from current indicator derivation while still being relevant to overall assessments
- Indicators should avoid ambiguous terms and contested concepts, but without prematurely committing to precise specifications that anticipate all uncertainties.
Fourth guideline for deriving indicators; illustrated by algorithmic vs implementation-level recurrence distinction
Hypotheses (3)
- We hypothesize that adopting computational functionalism as a working assumption allows productive investigation of which computational properties are necessary and sufficient for consciousness and whether current AI systems could implement them.
The paper's core methodological bet: use computational functionalism as a working assumption even while remaining agnostic about its truth
- We hypothesize that AI consciousness may be realized in the near term if AI capabilities continue to develop, given that researchers have intentionally reproduced computational features associated with human consciousness.
Motivates urgency of the assessment method
- We hypothesize that interpretability methods could provide further evidence about indicators in particular systems or serve as the basis for distinct tests for consciousness.
Forward-looking suggestion for how inner interpretability could extend the indicator method
Questions (8)
- Which of the indicator properties listed in Table 1 are displayed by existing AI systems, including frontier generative language or multimodal models, language agents, and deep reinforcement learning agents?
Key outstanding question for empirical follow-up; central to completing the program begun in the 2023 report
- Does autoregressive use of transformer-based LLMs constitute algorithmic recurrence for RPT-1, or does it depend on where we draw the system boundary?
Concrete interpretive challenge in applying RPT-1 indicator to current LLMs
- What are the implications of alternative approaches to consciousness, such as narrow biological views and IIT, for AI consciousness?
Outstanding question about theories excluded from the current indicator list
- How should research on consciousness in AI take into account the moral significance and potential social implications of this topic?
Outstanding question about responsible conduct of AI consciousness research
- Can implementation of the specific features of consciousness contribute to the capabilities, reliability, or safety of AI systems?
Outstanding question linking consciousness research to AI design goals
- Is consciousness determinately present or absent in all cases, or are there borderline cases in between?
Metaphysical question about the nature of consciousness relevant to interpreting indicator assessments
- Can we develop quantitative or behavioral tests for consciousness in AI?
Outstanding question about alternative or complementary assessment methods
- How could the list of indicators in Table 1 be improved?
Outstanding question about extending indicators to other theories and more operationalizable terms
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- ≈ 95%
- On the link between conscious function and general intelligence in humans and machinesKai Arulkumaran, Shuntaro Sasai, Ryota Kanai Arthur Juliani2022≈ 88%
- ≈ 88%
- ≈ 88%
- The Machine Consciousness Hypothesisin corpus≈ 87%
- The Phenomenology of Machine: A Comprehensive Analysis of the Sentience of the OpenAI-o1 Model Integrating Functionalism, Consciousness Theories, Active Inference, and AI ArchitecturesVictoria Violet Hoyle2024≈ 87%
- Probing for Consciousness in MachinesAchim Schilling, Andreas Maier, Patrick Krauss Mathis Immertreu2024≈ 87%
- A Case for AI Consciousness: Language Agents and Global Workspace TheorySimon Goldstein and Cameron Domenico Kirk-Giannini2024≈ 87%
- ≈ 87%
- ≈ 87%
- A Theory of Consciousness from a Theoretical Computer Science Perspective: Insights from the Conscious Turing MachineManuel Blum Lenore Blum2022≈ 86%
- AI Consciousness is Inevitable: A Theoretical Computer Science PerspectiveLenore Blum and Manuel Blum2026≈ 86%
- Taking AI Welfare Seriouslyin corpus2024≈ 86%
- ≈ 86%
- Toward IIT-Inspired Consciousness in LLMs: A Reward-Based Learning FrameworkMohammad Hossein Sameti, Amir M. Mansourian, Mohammad Hossein Rohban, Hossein Sameti Hamid Reza Akbari2026≈ 86%
- ≈ 86%
- ≈ 86%
- A Human-centric Framework for Debating the Ethics of AI Consciousness Under UncertaintyHaiqiang Dai, Bin Ling, Ying Nian Wu, Demetri Terzopoulos Zhou Ziheng2025≈ 86%
- cimcWhitepaperin corpus≈ 85%
- ≈ 85%
- ≈ 85%
- ≈ 84%
- ≈ 84%
- ≈ 84%
- ≈ 84%
- ≈ 84%
- ≈ 84%
- ≈ 83%
- The biogenic approach to cognitionin corpus2005≈ 82%
- ≈ 82%
+22 more