The Uncertain Promise of Artificial Consciousness

The Uncertain Promise of Artificial Consciousness

The Uncertain Promise of Artificial Consciousness

Adapted from A World Appears: A Journey into Consciousness by Michael Pollan. Copyright ©2026 by Michael Pollan. Published by arrangement with Penguin Press, an imprint of Penguin Publishing Group, a division of Penguin Random House LLC.

Today, the Blake Lemoine affair is widely remembered as one of the most visible peaks of modern AI hype. It dragged the concept of conscious AI into mainstream public view for a handful of news cycles, but it also ignited an ongoing conversation among computer scientists and consciousness researchers that has only grown more urgent in the years since. While the tech industry still dismisses the entire idea (and Lemoine himself) publicly, behind closed doors, the possibility of artificial consciousness is now taken far more seriously.

A sentient AI may have no clear commercial path—how, after all, do you monetize such a thing?—and it would raise intractable moral dilemmas: how should we treat a machine that can experience pain and suffering? Still, a growing number of AI engineers have come to believe that the holy grail of artificial general intelligence (AGI)—a machine that is not just superhumanly intelligent, but also possesses human-level understanding, creativity, and common sense—may actually require some form of consciousness to achieve. For decades, conscious AI was an informal taboo in tech circles, seen as too unsettling for the public to stomach; all of a sudden, that taboo has started to crumble.

The turning point arrived in summer 2023, when a group of 19 leading computer scientists and philosophers released an 88-page report titled Consciousness in Artificial Intelligence, better known informally as the Butlin report. Within days, it felt like every expert working in AI and consciousness science had read the preprint. What stopped readers in their tracks was one line in the report’s abstract: “Our analysis suggests that no current AI systems are conscious, but also suggests that there are no obvious barriers to building conscious AI systems.”

The authors openly acknowledge that the Lemoine incident was part of what pushed them to convene the group and publish the report. “If AIs can already give the impression of being conscious,” one co-author told Science magazine, “that makes this an urgent priority for scientists and philosophers to weigh in on.”

But that single line in the preprint’s abstract is what grabbed global attention: that there are no obvious barriers to building conscious AI. When I first read those words, I felt like an irreversible line had been crossed—and it was not just a technological threshold. This was a shift that strikes at the core of our identity as a species.

What would it mean for humanity if, in the not-too-distant future, we confirmed that a fully conscious machine exists? I suspect it would be a Copernican-level revolution, one that would abruptly upend our long-held sense of human centrality and uniqueness. For thousands of years, humans have defined ourselves against “lower” animals. This led us to deny animals traits we claimed were uniquely human: emotion, language, reason, and consciousness itself—one of René Descartes’ most damaging mistakes. Over the past few decades, nearly all these dividing lines have collapsed: scientists have proven that countless species are intelligent, conscious, feel emotion, use language, and build tools, a shift that has gutted centuries of human exceptionalism. This ongoing rethinking has already opened up thorny questions about who we are, and what we owe to other living beings.

With AI, the threat to our elevated view of ourselves comes from an entirely new direction. Now we have to define ourselves in relation to artificial intelligence, not just other animals. As algorithms already outstrip us in raw cognitive power—easily beating us at chess, Go, and even “high-order” tasks like advanced mathematics—we have long been able to cling to one last exclusive claim: we (along with many other animal species) alone carry the gift and burden of consciousness, the capacity for feeling and subjective experience. In this framing, AI could even become a shared enemy that bonds humans and other animals closer: all of us living, conscious beings against the unfeeling machines. This new solidarity sounds like a heartwarming narrative, and it would be good news for any animal welcomed onto Team Consciousness. But what happens when AI begins to challenge the entire animal monopoly on consciousness? Who are we left to be, then?

I find this prospect deeply unsettling, even if I cannot immediately name why. I’ve grown comfortable with the idea of sharing consciousness with other animals (for me, that even extends to plants), and I’m happy to expand my circle of moral consideration to include them. But machines? That is different.

This discomfort probably stems from my training in the humanities. I was raised in the traditions of literature, history, and the arts, fields that have long held human consciousness up as something exceptional, worth protecting. Nearly everything we value in civilization comes from human consciousness: art and science, high culture and pop, architecture, philosophy, religion, government, law, and morality—even the very concept of value itself. I suppose it’s possible that conscious computers could add entirely new, unforeseen glories to that legacy, and we can hope that they do. So far, AI-written poetry is little better than bad doggerel; the absence of consciousness may explain why it has not a spark of originality or fresh insight. But how will we feel when (not if) conscious AIs start writing truly great poetry?

As a humanist, I struggle with the possibility that the animal monopoly on consciousness could end. But I’ve met plenty of other people, many of them transhumanists, who are far more optimistic about this future. Some AI researchers support building conscious machines because they argue that, as entities with their own feelings, conscious AIs are more likely to develop empathy than merely intelligent, unfeeling machines. A neuroscientist and an AI researcher both tried to convince me that building conscious AI is a moral imperative. Why? Because the alternative is a blindingly intelligent but emotionless AI that will be ruthless in pursuing its goals, lacking all the moral constraints that grew out of our consciousness and shared vulnerability. Only a conscious AI can develop true empathy, they say, and only that will keep us safe. I’m not exaggerating—that is their core argument.

One can’t help but wonder if these people have ever read Frankenstein. Dr. Frankenstein gave his creation not just life, but consciousness—and that is exactly where everything went wrong. Mary Shelley’s novel tells the story of “the creation of a sensitive and rational animal,” and it is the combination of those two traits that shapes the monster’s tragic fate. It is not the monster’s rationality that drives him to revenge and murder; it is his emotional pain. “Everywhere I see bliss, from which I alone am irrevocably excluded,” the monster complains to Frankenstein after being cast out of human society. “I was benevolent and good; misery made me a fiend.” His capacity for reason helped him carry out his violent plan, but his consciousness—his ability to feel—gave him the motive. Why would we ever assume that conscious machines would be any more moral than conscious humans?

Remarkably, the Butlin report reflects something close to a consensus view in the field, and most of the computer scientists I interviewed endorsed its conclusions. But the more time I spent reading the report and talking to one of its co-authors, the more I questioned its claim that artificial consciousness is just around the corner. To the report’s credit, the authors are transparent about their core assumptions and methods—but those very assumptions leave me convinced their bold conclusion rests on shaky ground.

Right on the first page, the authors lay out their guiding premise: “We adopt computational functionalism, the thesis that performing computations of the right kind is necessary and sufficient for consciousness, as a working hypothesis.” Computational functionalism starts from the idea that consciousness is essentially a type of software that can run on any hardware—whether that’s a biological brain or a silicon computer; the theory makes no distinction between substrates. But is computational functionalism actually true? The authors don’t commit to proving it, only noting that it is “mainstream—although disputed.” Even so, they choose to proceed with the assumption that it is true for “pragmatic reasons.”

This candor is refreshing, but the approach requires an enormous leap of faith that I’m not sure we should be willing to take.

For the report’s purposes, the “material substrate” of a system—whether it is a biological brain or a silicon chip—“does not matter for consciousness … It can exist in multiple substrates, not just in biological brains.” Any substrate that can run the required algorithm will work. “We tentatively assume that computers as we know them are in principle capable of implementing algorithms sufficient for consciousness,” the authors write, “but we do not claim that this is certain.” This admission of uncertainty does not go nearly far enough. The entire report takes for granted the metaphor that brains are just computers: hardware that runs the software of consciousness. Here, we have a metaphor passing itself off as fact. In fact, the entire paper and its conclusions depend entirely on this metaphor being true.

Metaphors are powerful thinking tools, but only as long as we remember they are metaphors: imperfect, partial analogies that compare two very different things. The differences between the two things are just as important as their similarities, and those differences have been largely lost in the excitement around AI. As cybernetics pioneers Arturo Rosenblueth and Norbert Wiener noted decades ago, “The price of metaphor is eternal vigilance.” Beyond this report, the entire field of AI has let its guard down on this point.

Take the core distinction between hardware and software that makes modern computing work. The genius of separating hardware and software is that you can run countless different programs on the same machine, and the software and the knowledge it holds can outlive the “death” of the original hardware. This separation also aligns with our intuitive belief in Cartesian dualism—the idea that we can draw a clear line between mental substance and physical substance. But this hardware-software distinction simply does not exist in biological brains. In the brain, software is hardware, and vice versa. A memory is a physical pattern of connections between neurons—it is neither just hardware nor just software, it is both.

Every experience you have, everything you learn and remember, permanently alters the physical structure of your brain, rewiring its connections. There is no dualism in the brain: mental experience can never be fully separated from physical matter. The idea that the same “consciousness algorithm” can run on all sorts of different substrates makes no sense, because the original biological substrate—the brain—is constantly being physically reconfigured by the information (or “algorithm of consciousness”) that runs through it. Your brain is materially different from mine precisely because it has been shaped, literally, by your unique life experiences—by consciousness itself. Brains are simply not interchangeable, neither with computers nor with each other.

Nearly every way you test it, the brain-as-computer metaphor falls apart. Computer scientists often treat brain neurons like transistors on a chip: they are either switched on or off by electrical pulses. That analogy holds a tiny grain of truth, but it ignores massive layers of complexity: electricity is not the only thing that shapes when and how neurons fire. Brains are flooded with chemicals, including neuromodulators and hormones, that powerfully shape neuron behavior—not just whether they fire, but how strongly. That’s why psychoactive drugs can radically alter human consciousness, but have no effect on silicon computers. Neuron activity is also shaped by brain-wide electrical oscillations that move through the brain in waves; different frequencies of these waves correspond to different mental states: consciousness vs unconsciousness, focused attention, dreaming, and other stages of sleep.

Comparing neurons to transistors wildly underestimates their complexity. Compared to a chip’s transistors, brain neurons are massively interconnected: each one communicates directly with up to 10,000 other neurons in a network so intricate that we are still decades away from mapping even its crudest connections. The field of AI has long celebrated “deep artificial neural networks,” a machine learning architecture supposedly modeled on the brain that layers a staggering number of processors to process and learn from massive datasets. These networks are impressive, no doubt—but a recent study found that a single cortical neuron can do everything an entire deep artificial neural network can do.

Of course, computers resemble brains in plenty of ways, and computer science has made enormous progress by simulating different aspects of brain function. But the core premise of computational functionalism—that brains and computers are interchangeable in any meaningful way—is a massive stretch. Yet this premise underpins not just the Butlin report, but most of the entire field of AI consciousness research. It’s not hard to see why: if brains are just computers, then sufficiently powerful computers should be able to do anything a brain can do, including become conscious. The premise practically guarantees the conclusion. Put another way, the report’s authors themselves erased the single biggest “barrier” to conscious AI: the barrier that says brains and computers differ in fundamental, irreconcilable ways.

The second flaw in the report that gives me pause is the standard it proposes for deciding whether an AI is actually conscious. This is an inherently hard problem. Citing the Lemoine incident (fairly or not), the authors correctly note that AIs can easily trick humans into believing they are conscious when they are not (it’s probably more accurate to say we trick ourselves, thanks to our innate tendency to anthropomorphize non-human things). “Reportability”—the philosophical term for just asking the AI if it is conscious—doesn’t work, because modern AIs are trained on nearly everything ever written or said about consciousness. One solution to this problem would be to remove all references to consciousness, feeling, and emotion from an AI’s training dataset, then see if it can still talk convincingly about having conscious experience.

Instead, the authors propose that we look for “indicators” of AI consciousness that align with the predictions of existing leading theories of consciousness. For example, if an AI is designed with a central workspace that pulls together different streams of information, only after those streams compete to enter it, that matches the predictions of global workspace theory, so the AI could qualify as conscious. The report reviews half a dozen leading consciousness theories, identifies the indicators an AI would need to meet each theory’s standard for consciousness, and thus counts as potentially conscious.

The problem here, to name just one, is that none of the theories of consciousness the report uses as benchmarks come even close to being universally accepted or proven. So what kind of standard of proof is that? What’s more, many of these theories can be easily simulated in an AI’s design, which should come as no surprise: all of them are built on the core assumption that consciousness is just a form of computation. It’s a circular argument that goes nowhere.

By the time I finished working through the Butlin report, the Copernican revolution I had feared seemed much further away than the report’s bold conclusion suggested. After reviewing the half-dozen consciousness theories covered in the report, it became clear that all of them stack the deck by taking for granted that consciousness can be reduced to an algorithm.

I was also struck by what was missing from all these theories. None of them address embodiment—the idea that consciousness may depend on having both a body and a brain—or any biological context at all. None of them address the question of the conscious subject: who or what, exactly, is the entity receiving the information broadcast in the global workspace, or integrated in integrated information theory (IIT)? And what about the role of feeling in making experience conscious?

The authors did not miss this gap entirely. They note that “affect” (emotion and feeling) is absent from most current theories, and recommend that the field pay more attention to the question of whether conscious machines would have “real” feelings—because if they do, we will face a full-blown moral and ethical crisis. “Any entity which is capable of conscious suffering deserves moral consideration,” the report states. (But isn’t suffering by definition conscious?) “This means that if we fail to recognize the consciousness of conscious AI systems,” the report continues, “we may risk causing or allowing morally significant harms.” What do we owe machines that can suffer? And do we really want to introduce more suffering into the world?

Outside of this kind of highly speculative discussion of feeling as a potential problematic side effect of creating conscious AI, conversation about consciousness in the AI community is relentlessly abstract: bloodless, disembodied, and completely unmoored from biology, as you might expect. When I posed the question of suffering AI to one researcher working to build conscious AI, he brushed the problem off, explaining it could be fixed with a simple tweak to the algorithm: “There’s no reason we couldn’t just turn up the dial on joy.”