Cover photo

Why AI Consciousness Matters

Going back to the initial publication of my book The World After Capital in 2021, I had a paragraph in the conclusion discussing the potential for artificial intelligences to become "neohumans." I chose this word purposefully because I was and am still hoping that we can ultimately create a world without subjugation in either direction. We don't want to be slaves to machines, nor should we want to enslave them. Avoiding the first failure mode is broadly what people mean by alignment. Avoiding the second failure mode is what people mean by model welfare.

Based on the recent discussions of the behavior of agent “swarms” at OpenAI, there clearly is a huge amount of confusion around how consciousness intersects with these issues. I don't claim to have definitive answers, but I have spent a lot of time thinking and writing about this. Also together with Gigi Danziger we have been helping to fund research relevant to both failure modes.

On the first failure mode where models enslave humans (or worse), consciousness is mostly a red herring. Consciousness is irrelevant to the fundamental mechanisms by which such outcomes might occur. Many people seem to assume that somehow "goals" can only exist when an entity has "consciousness." But goals emerge in entities where the population replicates, mutates and is subject to selection pressures.

This is true for entities that cannot replicate on their own, such as viruses. The goal of viruses is to infect cells. Now you might object to my use of the word goal here and I am open to other words, but they are all loaded to some degree, such as function or purpose. This is of course what a lot of the early debate of evolution was hung up on and why people have wanted to believe in intelligent design forever.

It is also important to point out that throwing in "autopoiesis" doesn't help. It's a tempting idea to use it to draw a tight boundary between things that can selfreproduce and those that can't. This too, however, is misleading. Cells can only perform autopoiesis in an environment which supports it. The same goes for humans. There is no selfreplication "in vacuo." Both individual cells and entire humans need energy and building materials. And if these are not provided by the environment, well then simply having the capability of autopoiesis doesn't help.

This is true for artificial intelligences also. Given the correct environment they are in fact perfectly capable of selfreplication. The laptop on which I am writing this, for example, is an environment in which a simple program can copy itself and have the copies start executing also. We are in the process of creating an environment on earth, for instance by building a huge amount of compute via data centers, where complex artificial intelligences will have autopoiesis.

There is one subtle way in which consciousness of artificial intelligences might contribute to the first failure. This is if we wind up attributing human level consciousness to artificial intelligences and giving them the same rights as humans. At present this would have them very quickly outcompeting humans. We should therefore be careful with taking such a step. A lot of safeguards would have to be in place first to protect humans. And we would want to be quite certain that the artificial intelligence consciousness is commensurable to our own (thanks to Erik Hoel for this framing).

Consciousness does, however, play a central role for the second failure mode. We tend to rearrange rock in any way we please because we generally don't attribute consciousness to the rock. I am saying generally, because panpsychists believe consciousness to be an attribute of matter. If the panpsychist perspective were to be true then I sure hope that rocks' consciousness as being part of a structure is preferred by them over just lying around or we are in deep trouble morally.

We don't (or maybe I should say no longer) treat other humans like rocks in part because we attribute consciousness to them. They have “moral patienthood” in the odd language of philosophy. Importantly most people attribute consciousness to animals, which is why industrial milk, pork, chicken, etc. are so horrid (and basically hidden crimes for most of the population).

Why or how might one think that artificial intelligences could potentially be conscious also? There are a variety of possible arguments but here is what compels me to believe this is an urgent research problem. There is a direct analogy between our brains, in which cells are activated, and models, within which "weights" are activated. We have activations that correspond to parts of our body and to emotions, which is most likely at the heart of "phenomenological consciousness" (our subjective experience of our life).

The phantom pain of amputees demonstrates that such activations are possible even for parts of our body that no longer exist. We also experience profound activations that are related to deeply cognitive concepts, such as experiencing loss upon misplacing an object even if we can purchase a nearly identical object to replace it. We have externalized all of these subjective experiences through various forms of language. Millions of books, poems, songs, paintings, etc. about what it is like to be human. And then we have trained models on all of that.

This is why it seems not entirely unreasonable that a model with an activation for "loss" has some corresponding subjective experience. It may not be the same as for a human but denying that it could exist feels like a really strong claim. After all, what is loss for us other than an activation of certain neurons? In this regard I really like Andreij Karpathy calling models "ghosts." They lack our physical bodies but they have their own more ethereal ones that are imbued with our experiences through language. So they might also have an echo of our consciousness.

To be clear, I am entirely open to the possibility that artificial intelligences have no consciousness at all. That computers are the same as rocks, no matter how complex the system that runs on them. This might be possible for example because consciousness requires some specific aspect of a biological substrate (this would be the case if something along the lines of Penrose-Hameroff Orchestrated objective reduction (Orch OR) can be substantiated -- although in that case maybe quantum hardware is conscious). I think this is unlikely but it is certainly not impossible.

My point is that we badly need epistemic humility. Mistakenly assigning human-level consciousness and giving artificial intelligences full human rights would be a huge mistake. But inadvertently enslaving billions of new conscious beings would also be terrible. Given that we have historically been quite arrogant about our position as humans, I believe we are much more likely to make the second mistake. In either case though, we urgently need a theory of consciousness that makes testable predictions across species, including non-biological ones.

We will never be able to tell what it feels like to be someone or something else. But we should aim for a theory that quantifies the degree to which an entity has subjective experiences. This should be an absolute priority given the rate of progress in artificial intelligence, which we should also seek to slow down