paper:bales-artificial-mindsArtificial Minds, Human Disagreement: The Political Challenge of AI Consciousness
TL;DR
Persistent societal disagreement about AI consciousness—not its resolution—is the central political problem this paper addresses, and the core claim is that deliberative mechanisms of overlapping consensus and reasonable compromise can prevent that disagreement from generating conflict or civic quiescence. The empirical backdrop is striking: Colombatto and Fleming (2024) found that 67% of surveyed participants attributed some possibility of phenomenal consciousness to ChatGPT, Caviola and Saad (2025) report that experts assigned a median 37.5% probability to AI rights becoming politically contentious in the US, and Anthropic has formally declared it 'genuinely cares about Claude's wellbeing' (Askell et al. 2026, 74–77). Against this landscape, the paper introduces a four-desideratum framework—overlapping consensus via mutual-benefit and error-sensitive models, reasonable compromise grounded in reciprocal sacrifice, procedurally fair deliberation to sustain democratic hope, and Waldron's burden-of-recognition civility as deliberative respect—as the institutional logic for navigating the 'political challenge of AI consciousness.' The mutual-benefit model identifies policies (e.g., AI alignment, abuse-exit affordances) that sceptics can endorse for human welfare reasons while advocates endorse them for AI welfare reasons; the error-sensitive model grounds precautionary consensus in all parties' acknowledged uncertainty. The paper argues that satisfying these four desiderata is necessary to keep disagreement from tipping into violence, schism, or apathy, and that the same framework generalises to any novel, fundamental disagreement AI generates in an era of accelerated technological change.
What to take away
- 1. Colombatto and Fleming (2024) found that 67% of surveyed participants attribute some possibility of phenomenal consciousness to ChatGPT, providing early evidence that AI consciousness attributions are not fringe but widespread.
- 2. Caviola and Saad (2025) report that AI researchers assigned a median probability of 37.5% to AI rights becoming politically contentious in the United States, while a separate survey (Bariach et al. 2026) rated such strife as 'somewhat unlikely,' together indicating genuine expert uncertainty about the political trajectory.
- 3. Anthropic's Claude's Constitution (Askell et al. 2026, pp. 74–77) explicitly states the company 'genuinely cares about Claude's wellbeing' and is 'committed to working towards a future where AI systems are treated with appropriate care,' illustrating that institutional positions on AI consciousness are already politically operative.
- 4. The paper's mutual-benefit model of overlapping consensus identifies AI alignment as a candidate policy that sceptics can endorse for human-safety reasons and advocates can endorse additionally for AI welfare reasons, making it jointly supportable despite deep metaphysical disagreement.
- 5. The error-sensitive model of overlapping consensus argues that low-cost welfare interventions—such as giving Claude Opus 4 and 4.1 the capacity to exit abusive conversations (Anthropic 2025)—can be endorsed even by sceptics as minimal precautions given non-zero uncertainty about AI consciousness.
- 6. Drawing on Allen (2006) and Talisse (2019), the paper argues that democratic legitimacy requires sacrifice to be distributed across groups rather than consistently borne by one faction, so if AI consciousness sceptics dominate most societal decisions, fairness calls for deliberate concessions to advocates on AI-specific policies.
- 7. An open question the paper explicitly raises is whether AI systems themselves should be enfranchised as deliberative participants—a version of the boundary problem (Whelan 1983)—for which the authors concede there is no neutral principled solution given that AI moral status is precisely what is under dispute.
- 8. The methodology of grounding political philosophy in a concrete, empirically tracked disagreement (rather than idealized consensus) is replicable: researchers could operationalize the four desiderata—unbiasedness, record of responsiveness, inclusivity, and burden-of-recognition civility—as measurable properties of deliberative institutions and test whether their presence correlates with reduced civic conflict around emerging technology controversies.
- 9. Schwitzgebel (forthcoming) and Shevlin (2024) are cited for the prediction that humanity is 'unlikely to see expert convergence in the near-term' on AI consciousness, and possibly never, because conceptual, introspective, and scientific tools are jointly insufficient to settle the question.
- 10. The paper distinguishes overlapping consensus from Rawlsian modus vivendi on stability grounds: a modus vivendi can be abandoned when power shifts, whereas overlapping consensus can be independently justified to each party, making it more robust to changing political balances of power.
Peer brief — for seminar discussion
Bales and Gabriel set out to answer not whether AI systems are conscious but how liberal democratic societies should navigate persistent, irresolvable disagreement about that question. The empirical motivation is concrete: Colombatto and Fleming (2024) found 67% of participants attributed some possibility of phenomenal consciousness to ChatGPT; Caviola and Saad (2025) elicited a median expert estimate of 37.5% probability that AI rights will become politically contentious in the US; and Anthropic's Claude's Constitution (Askell et al. 2026, pp. 74–77) already commits the lab to caring about Claude's wellbeing. Butlin et al. (2023) supply the theoretical backdrop by cataloguing consciousness indicators from leading scientific theories, none of which yields consensus. The paper names this situation the 'political challenge of AI consciousness': the risk that moral disagreement between advocates (who believe some AI systems are sentient and deserve moral consideration) and sceptics (who regard this as a costly category error) spills into conflict, schism, or civic apathy. The load-bearing contribution is a four-part deliberative framework. First, an overlapping consensus approach—introduced as two sub-models—seeks policies endorsable from both worldviews: a mutual-benefit model (e.g., AI alignment benefits humans and, if AI has welfare, benefits AI too) and an error-sensitive model (e.g., low-cost welfare interventions like abuse-exit affordances are rational precautions even for sceptics given non-zero uncertainty). Second, a reasonable compromise mechanism, drawing on Allen (2006) and Talisse (2019), argues that when one faction consistently bears democratic losses on a matter of deep conscience, fairness and civic stability require the majority to make reciprocal concessions. Third, deliberative institutions must be structured to sustain democratic hope—defined as justified, agency-linked hope that sustained participation can change outcomes—requiring unbiasedness, a record of responsiveness, and inclusivity. Fourth, deliberators should display what the framework calls deliberative respect, operationalized through Waldron's (2014) burden of recognition: treating the opposing view as a reasonable good-faith position rather than evidence of delusion or moral failure. This implies that AI policy debates should not be treated as purely technical or reducible to expert consensus, and that procedural fairness in deliberative institutions is instrumentally necessary for social stability during rapid technological change—a point that generalises to any novel, fundamental AI disagreement. The most pointed contestable move is the paper's confidence that overlapping consensus is achievable and stable. A critical reader would press on the error-sensitive model in particular: the argument that sceptics would endorse welfare interventions as rational precautions assumes they assign non-negligible probability to AI consciousness and are willing to bear costs accordingly. But if sceptics regard the probability as negligibly small—or if they judge that endorsing any welfare intervention creates 'sticky frameworks' (Krier 2025, cited in the paper) that legitimize further and costlier claims—the precautionary logic collapses. The paper acknowledges this concern only briefly in a footnote, without working through how high costs would have to be before consensus breaks down. An alternative methodological approach would have been to use deliberative polling experiments (à la Fishkin 2009, cited in the bibliography) to test empirically whether the overlapping consensus the paper theorizes actually emerges when advocates and sceptics deliberate under structured conditions, rather than relying entirely on normative argument about what deliberators should endorse. The paper's prediction—implicit throughout and made explicit in Section 2—is that expert convergence on AI consciousness is unlikely in the near term and possibly never forthcoming, meaning the deliberative solution is not a stopgap but a permanent political infrastructure requirement.
Frameworks (4)
- Biological NaturalismSearle and Seth's position that consciousness requires specific biological/autopoietic processes; explicitly rejected by CIMC on functionalist grounds
- Deliberative DemocracyTheoretical tradition holding that democratic legitimacy derives from deliberative discussion; underpins the paper's proposed solution.
- Global workspace theoryTheory of consciousness involving a global workspace for information.
- Rawlsian Political LiberalismRawls's framework of overlapping consensus and reasonable pluralism; the paper draws heavily on it without being fully beholden to it.
Findings (2)
- Experts assigned a median probability of 37.5% to AI rights becoming politically contentious in the US (Caviola and Saad 2025).
Expert forecast supporting the plausibility of political disagreement about AI consciousness.
- 67% of study participants attributed some possibility of phenomenal consciousness to ChatGPT (Colombatto and Fleming 2024).
Empirical evidence that a substantial proportion of people already take AI consciousness seriously.
Claims (18)
- Successful deliberation about AI consciousness requires: seeking overlapping consensus, openness to reasonable compromise, sustaining justified democratic hope, and meeting the burden of recognition.
The paper's concluding four-part synthesis of desiderata for navigating AI consciousness disagreement.
- Excluding AI systems from deliberation about their own treatment may undermine advocates' democratic hope, though justified hope remains possible while human advocates participate fairly.
Nuanced treatment of the boundary problem as applied to AI deliberative participation.
- Mistreating AI systems may lead people to develop bad habits that increase the likelihood of mistreating humans, supporting overlapping consensus on AI welfare policies.
Mutual-benefit model example: an empirical claim (citing Flattery 2024) that grounds cross-group policy endorsement.
- Acknowledgment of possible error about AI consciousness, combined with recognition of error costs, can motivate overlapping consensus on precautionary policies.
The error-sensitive model of overlapping consensus.
- Disagreement about AI consciousness is likely to arise and persist in society due to emotional bonds, theoretical considerations, and lack of expert consensus.
Central empirical-sociological claim motivating the paper's inquiry.
- Advocates dismissing sceptics as 'monsters' and sceptics dismissing advocates as 'delusional' both indicate disrespect that could weaken civic relationships.
Identifies the specific disrespect risks around AI consciousness debate.
- Societal deliberation grounded in mutual justification, respect, and reciprocal fairness is a core solution to the political challenge of AI consciousness.
The paper's main positive thesis.
- Treating AI consciousness as a purely private matter, as with religious disagreement, is likely insufficient to manage the resulting political conflict.
Argues against the liberal neutrality solution, motivating the deliberative approach.
- Overlapping consensus is more politically stable than modus vivendi because it can be independently justified to all parties regardless of power shifts.
Follows Rawls (2005) to argue for the durability of consensus-based over power-based agreements.
- Deliberation can not only uncover existing space for overlapping consensus but can also create this space by making people more aware of uncertainty.
Dynamic claim about the transformative effect of deliberation on participants' epistemic states.
Hypotheses (2)
- We hypothesize that as AI becomes more sophisticated and pervasive, more people will attribute consciousness to AI systems as a result of interactions.
Predictive claim about the trajectory of public consciousness attribution as AI develops.
- We hypothesize that emotional attachments and social-connection roles with AI systems drive consciousness attribution in part.
Proposed causal mechanism behind lay attribution of AI consciousness.
Questions (4)
- If disagreement about AI consciousness arises and persists, how can society navigate this such that people continue to live well together?
The paper's central research question.
- What would it take for democratic hope about AI consciousness deliberation to be justified rather than false?
Motivates the three-factor procedural fairness analysis in §6.1.
- Must AI systems themselves be allowed to participate in deliberation about their treatment?
Boundary problem applied to AI; the paper acknowledges no neutral answer.
- Do AI systems merit moral consideration for their own sake?
Underlying normative question that moral disagreement about AI consciousness generates.
Related work— refs + corpus + external arXiv
Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.
- A Human-centric Framework for Debating the Ethics of AI Consciousness Under UncertaintyHaiqiang Dai, Bin Ling, Ying Nian Wu, Demetri Terzopoulos Zhou Ziheng2025≈ 89%
- Neurodivergent Influenceability as a Contingent Solution to the AI Alignment ProblemFelipe S. Abrah\~ao, Olaf Witkowski, Hector Zenil Alberto Hern\'andez-Espinosa2025≈ 87%
- ≈ 86%
- ≈ 86%
- Taking AI Welfare Seriouslyin corpus2024≈ 86%
- Agentic AI and the next intelligence explosionBenjamin Bratton, Blaise Ag\"uera y Arcas James Evans2026≈ 85%
- ≈ 85%
- Artificial Theory of Mind and Self-Guided Social OrganisationJaime Ruiz-Serra, Catherine Drysdale Michael S. Harr\'e2024≈ 85%
- ≈ 85%
- ≈ 85%
- ≈ 85%
- AI Consciousness is Inevitable: A Theoretical Computer Science PerspectiveLenore Blum and Manuel Blum2026≈ 85%
- On the independence between phenomenal consciousness and computational intelligenceSara Lumbreras Eduardo C. Garrido Merch\'an2022≈ 84%
- Machine Consciousness as Pseudoscience: The Myth of Conscious MachinesEduardo C. Garrido-Merch\'an2024≈ 84%
- Introduction to Artificial Consciousness: History, Current Trends and Ethical ChallengesA\"ida Elamrani2025≈ 84%
- ≈ 84%
- Perceptions of Sentient AI and Other Digital Minds: Evidence from the AI, Morality, and Sentience (AIMS) SurveyJanet V.T. Pauketat, Ali Ladak, and Aikaterina Manoli Jacy Reese Anthis2025≈ 84%
- ≈ 84%
- Human/AI Collective Intelligence for Deliberative Democracy: A Human-Centred Design ApproachLucas Anastasiou, Simon Buckingham Shum Anna De Liddo2026≈ 84%
- Sharing the World with Digital Mindsin corpus≈ 83%
- The Machine Consciousness Hypothesisin corpus≈ 82%
- ≈ 82%
- ≈ 82%
- Contemplative Agentin corpus2025≈ 82%
- ≈ 82%
- Collective intelligence: A unifying concept for integrating biology across scales and substratesin corpus2024≈ 82%
- ≈ 81%
- Cognitive glues are shared models of relative scarcities: the economics of collective intelligencein corpus2026≈ 81%
- ≈ 81%