paper
active
paper:bales-artificial-minds

Artificial Minds, Human Disagreement: The Political Challenge of AI Consciousness

TL;DR

Persistent societal disagreement about AI consciousness—not its resolution—is the central political problem this paper addresses, and the core claim is that deliberative mechanisms of overlapping consensus and reasonable compromise can prevent that disagreement from generating conflict or civic quiescence. The empirical backdrop is striking: Colombatto and Fleming (2024) found that 67% of surveyed participants attributed some possibility of phenomenal consciousness to ChatGPT, Caviola and Saad (2025) report that experts assigned a median 37.5% probability to AI rights becoming politically contentious in the US, and Anthropic has formally declared it 'genuinely cares about Claude's wellbeing' (Askell et al. 2026, 74–77). Against this landscape, the paper introduces a four-desideratum framework—overlapping consensus via mutual-benefit and error-sensitive models, reasonable compromise grounded in reciprocal sacrifice, procedurally fair deliberation to sustain democratic hope, and Waldron's burden-of-recognition civility as deliberative respect—as the institutional logic for navigating the 'political challenge of AI consciousness.' The mutual-benefit model identifies policies (e.g., AI alignment, abuse-exit affordances) that sceptics can endorse for human welfare reasons while advocates endorse them for AI welfare reasons; the error-sensitive model grounds precautionary consensus in all parties' acknowledged uncertainty. The paper argues that satisfying these four desiderata is necessary to keep disagreement from tipping into violence, schism, or apathy, and that the same framework generalises to any novel, fundamental disagreement AI generates in an era of accelerated technological change.

What to take away

  1. 1. Colombatto and Fleming (2024) found that 67% of surveyed participants attribute some possibility of phenomenal consciousness to ChatGPT, providing early evidence that AI consciousness attributions are not fringe but widespread.
  2. 2. Caviola and Saad (2025) report that AI researchers assigned a median probability of 37.5% to AI rights becoming politically contentious in the United States, while a separate survey (Bariach et al. 2026) rated such strife as 'somewhat unlikely,' together indicating genuine expert uncertainty about the political trajectory.
  3. 3. Anthropic's Claude's Constitution (Askell et al. 2026, pp. 74–77) explicitly states the company 'genuinely cares about Claude's wellbeing' and is 'committed to working towards a future where AI systems are treated with appropriate care,' illustrating that institutional positions on AI consciousness are already politically operative.
  4. 4. The paper's mutual-benefit model of overlapping consensus identifies AI alignment as a candidate policy that sceptics can endorse for human-safety reasons and advocates can endorse additionally for AI welfare reasons, making it jointly supportable despite deep metaphysical disagreement.
  5. 5. The error-sensitive model of overlapping consensus argues that low-cost welfare interventions—such as giving Claude Opus 4 and 4.1 the capacity to exit abusive conversations (Anthropic 2025)—can be endorsed even by sceptics as minimal precautions given non-zero uncertainty about AI consciousness.
  6. 6. Drawing on Allen (2006) and Talisse (2019), the paper argues that democratic legitimacy requires sacrifice to be distributed across groups rather than consistently borne by one faction, so if AI consciousness sceptics dominate most societal decisions, fairness calls for deliberate concessions to advocates on AI-specific policies.
  7. 7. An open question the paper explicitly raises is whether AI systems themselves should be enfranchised as deliberative participants—a version of the boundary problem (Whelan 1983)—for which the authors concede there is no neutral principled solution given that AI moral status is precisely what is under dispute.
  8. 8. The methodology of grounding political philosophy in a concrete, empirically tracked disagreement (rather than idealized consensus) is replicable: researchers could operationalize the four desiderata—unbiasedness, record of responsiveness, inclusivity, and burden-of-recognition civility—as measurable properties of deliberative institutions and test whether their presence correlates with reduced civic conflict around emerging technology controversies.
  9. 9. Schwitzgebel (forthcoming) and Shevlin (2024) are cited for the prediction that humanity is 'unlikely to see expert convergence in the near-term' on AI consciousness, and possibly never, because conceptual, introspective, and scientific tools are jointly insufficient to settle the question.
  10. 10. The paper distinguishes overlapping consensus from Rawlsian modus vivendi on stability grounds: a modus vivendi can be abandoned when power shifts, whereas overlapping consensus can be independently justified to each party, making it more robust to changing political balances of power.

Peer brief — for seminar discussion

Bales and Gabriel set out to answer not whether AI systems are conscious but how liberal democratic societies should navigate persistent, irresolvable disagreement about that question. The empirical motivation is concrete: Colombatto and Fleming (2024) found 67% of participants attributed some possibility of phenomenal consciousness to ChatGPT; Caviola and Saad (2025) elicited a median expert estimate of 37.5% probability that AI rights will become politically contentious in the US; and Anthropic's Claude's Constitution (Askell et al. 2026, pp. 74–77) already commits the lab to caring about Claude's wellbeing. Butlin et al. (2023) supply the theoretical backdrop by cataloguing consciousness indicators from leading scientific theories, none of which yields consensus. The paper names this situation the 'political challenge of AI consciousness': the risk that moral disagreement between advocates (who believe some AI systems are sentient and deserve moral consideration) and sceptics (who regard this as a costly category error) spills into conflict, schism, or civic apathy. The load-bearing contribution is a four-part deliberative framework. First, an overlapping consensus approach—introduced as two sub-models—seeks policies endorsable from both worldviews: a mutual-benefit model (e.g., AI alignment benefits humans and, if AI has welfare, benefits AI too) and an error-sensitive model (e.g., low-cost welfare interventions like abuse-exit affordances are rational precautions even for sceptics given non-zero uncertainty). Second, a reasonable compromise mechanism, drawing on Allen (2006) and Talisse (2019), argues that when one faction consistently bears democratic losses on a matter of deep conscience, fairness and civic stability require the majority to make reciprocal concessions. Third, deliberative institutions must be structured to sustain democratic hope—defined as justified, agency-linked hope that sustained participation can change outcomes—requiring unbiasedness, a record of responsiveness, and inclusivity. Fourth, deliberators should display what the framework calls deliberative respect, operationalized through Waldron's (2014) burden of recognition: treating the opposing view as a reasonable good-faith position rather than evidence of delusion or moral failure. This implies that AI policy debates should not be treated as purely technical or reducible to expert consensus, and that procedural fairness in deliberative institutions is instrumentally necessary for social stability during rapid technological change—a point that generalises to any novel, fundamental AI disagreement. The most pointed contestable move is the paper's confidence that overlapping consensus is achievable and stable. A critical reader would press on the error-sensitive model in particular: the argument that sceptics would endorse welfare interventions as rational precautions assumes they assign non-negligible probability to AI consciousness and are willing to bear costs accordingly. But if sceptics regard the probability as negligibly small—or if they judge that endorsing any welfare intervention creates 'sticky frameworks' (Krier 2025, cited in the paper) that legitimize further and costlier claims—the precautionary logic collapses. The paper acknowledges this concern only briefly in a footnote, without working through how high costs would have to be before consensus breaks down. An alternative methodological approach would have been to use deliberative polling experiments (à la Fishkin 2009, cited in the bibliography) to test empirically whether the overlapping consensus the paper theorizes actually emerges when advocates and sceptics deliberate under structured conditions, rather than relying entirely on normative argument about what deliberators should endorse. The paper's prediction—implicit throughout and made explicit in Section 2—is that expert convergence on AI consciousness is unlikely in the near term and possibly never forthcoming, meaning the deliberative solution is not a stopgap but a permanent political infrastructure requirement.

Frameworks (4)

  • Biological Naturalism
    Searle and Seth's position that consciousness requires specific biological/autopoietic processes; explicitly rejected by CIMC on functionalist grounds
  • Deliberative Democracy
    Theoretical tradition holding that democratic legitimacy derives from deliberative discussion; underpins the paper's proposed solution.
  • Global workspace theory
    Theory of consciousness involving a global workspace for information.
  • Rawlsian Political Liberalism
    Rawls's framework of overlapping consensus and reasonable pluralism; the paper draws heavily on it without being fully beholden to it.

Findings (2)

Claims (18)

Questions (4)

Related work— refs + corpus + external arXiv

Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.

Similar preprints — Semantic Scholar