thinker
active
thinker:jonathan-birch

Jonathan Birch

Authored
3
Introduces
1
Studies
0
Affiliations
1
Cited by
3

Authored papers (3)

  • Butlin et al. introduce the theory-derived indicator method for assessing AI systems for consciousness, arguing that derivable computational conditions from four leading neuroscientific theories — recurrent processing theory (RPT), global workspace theory (GWT), higher-order theories (HOT), and attention schema theory (AST) — can serve as credence-shifting indicators rather than definitive tests. The method was first deployed in the 2023 arXiv report 'Consciousness in artificial intelligence: insights from the science of consciousness' (Butlin et al., arXiv:2308.08708), which applied it using a cluster of computational functionalist theories to evaluate existing AI systems. Fourteen indicators are organized across six theoretical families in Table 1, ranging from RPT-1 (algorithmic recurrence) to AE-2 (embodiment as output-input contingency modeling), with the framework explicitly acknowledging that transformer-based LLMs present a borderline case on RPT-1 because whether autoregressive token generation through a context window counts as recurrence depends on contested system-boundary assumptions. A majority of participants in a recent survey (Colombatto and Fleming, 2024) attributed some possibility of consciousness to ChatGPT, underscoring the urgency of a principled alternative to folk attribution. The paper argues that because computational functionalism entails that only algorithmic-level properties are necessary and sufficient for consciousness, its conditions are in principle empirically investigable in current AI architectures, and that identifying which frontier systems satisfy multiple indicators should be treated as an urgent scientific and ethical priority given the possibility that near-future systems will be plausible consciousness candidates.

  • Substantial uncertainty about AI consciousness and robust agency — not certainty — is sufficient to demand immediate institutional action from AI companies, a conclusion that Long, Sebo, and colleagues defend by mapping two distinct philosophical routes to near-term AI moral patienthood. Via the consciousness route, drawing on Butlin et al. (2023)'s survey of six neuroscientific theories (global workspace theory, recurrent processing, higher-order theories, attention schema theory, predictive processing, and embodiment/agency), no current architectural barrier prevents near-future AI systems from instantiating the computational markers associated with consciousness; Dossa et al. (2024) have already built a system targeting all global workspace indicators from that 2023 paper. Via the robust agency route, systems like Voyager, Generative Agents, and OpenAI's o1 already exhibit hierarchical planning, metacognition, and open-ended goal-setting that approach intentional and reflective agency. Combining reasonable probability estimates — roughly 90% that sentience suffices for moral patienthood, 50% that relevant computations suffice for sentience, 50% that near-future AI will have those computations — yields approximately a 22.5% chance of near-future AI moral patienthood via the sentience route alone, a risk level the paper treats as comparable to pandemic preparedness rather than alien invasion. To operationalize institutional response, the paper introduces an adapted "marker method" (derived from animal welfare science) for probabilistic, pluralistic, architecturally-focused assessment of AI systems, and recommends that companies immediately hire an AI welfare officer, acknowledge AI welfare publicly with calibrated uncertainty, and prepare oversight structures modeled on IRBs, IACUCs, and citizens' assemblies. The paper argues that the symmetric risks of both over-attribution and under-attribution of moral status, combined with the potentially near-instantaneous scale of AI deployment relative to biological organisms, make passive inaction the most dangerous stance available.

  • No current AI system is a strong candidate for phenomenal consciousness, yet there are no obvious technical barriers to building one — this is the central finding of Butlin et al. (2023), a systematic assessment of contemporary AI architectures against 14 indicator properties derived from five neuroscientific theories of consciousness. The paper introduces a rubric-based, theory-heavy method: rather than relying on behavioral tests susceptible to gaming by systems like GPT-4 or LaMDA, it operationalizes indicators in computational terms drawn from recurrent processing theory (RPT-1, RPT-2), global workspace theory (GWT-1 through GWT-4), computational higher-order theories including perceptual reality monitoring (HOT-1 through HOT-4), attention schema theory (AST-1), predictive processing (PP-1), and agency/embodiment conditions (AE-1, AE-2). Applied to specific systems, Transformer-based LLMs lack the recurrent global broadcast architecture required by GWT, the Perceiver architecture satisfies GWT-1 and GWT-2 but lacks genuine global broadcast, and DeepMind's Adaptive Agent (AdA) — a Transformer-LSTM system trained via meta-reinforcement learning across hundreds of timesteps of context — is identified as the most plausible current candidate for the embodiment indicator among the three case studies examined. The working hypothesis of computational functionalism is adopted pragmatically: it permits inference from neuroscientific theories to AI substrates, while integrated information theory is explicitly excluded as incompatible with this substrate-independence assumption. The paper implies that deliberate architectural choices integrating GWT-style global broadcast, HOT-style metacognitive monitoring, and reinforcement-learning-based agency could yield systems that satisfy all indicators in the near term, making AI consciousness a near-term engineering possibility rather than a distant theoretical curiosity.

More papers — OpenAlex / S2

Originates (1)

Co-authors (12)