Preventing Apocalypse or a Political Power Grab? Speaking with AI Safety Expert David Krueger

Absent a change of course on AI safety by governments worldwide, former Cambridge professor and leading AI-safety expert David Krueger believes that human extinction is "more likely than not." He makes the case to our own Henri Mattila, a skeptic of both doomsday scenarios and political overreach.

In early 2024, David Krueger left the organization he had spent six months helping to establish: the United Kingdom’s AI Safety Institute.

The American computer scientist was among a growing number of researchers who believed that the technology behind the hugely popular app ChatGPT might prove to be a wolf in sheep’s clothing. In his view, large language models and other rapidly advancing forms of artificial intelligence (AI) pose an under-appreciated threat to civilization itself, even one of apocalyptic proportions. Yet most governments, including that of the United States, have tended to appear more concerned with maintaining their lead in the race for AI technical capability.

Krueger found a more receptive audience in the United Kingdom, where he helped to establish one of the world’s first government institutes devoted to studying the risks posed by advanced AI and advising officials on how to prepare for worst-case scenarios. Since leaving, however, Krueger has suggested that many of the methods are inadequate to the scale of the challenge, describing them as "silly" and "BS."

He is now among the most vocal advocates of a far more muscular and internationally coordinated response.

Through his new organization, Evitable, Krueger wants to “put an end to the reckless race to develop superintelligence,” with enforcement resting primarily with governments—above all, those of the United States and China. He has called for “enforceable, international limits on AI capabilities” and suggested that the most robust way to stop the race would be to cease producing advanced AI chips.

Such proposals can sound reasonable enough. With that said, their implications become more troubling once one considers what enforcement would require.

Behind the technical language, such proposals would effectively grant governments extraordinary authority over information and private enterprise. In the United States, such controls would cut against the spirit—and perhaps the letter—of the First Amendment, as well as that of the free market. AI systems are, at bottom, code and information, ones and zeros. And if they prove capable of substantially increasing economic productivity, society at large stands to benefit. When the government introduces a heavy-handed regulatory regime, though, innovation tends to stagnate, the most politically connected firms concentrate power among themselves, and the rest of us are left behind.

After speaking with Krueger over the course of two interviews this spring and early summer, however, I found that he had considered many such counterarguments.

To Krueger and other activists concerned with existential AI risk, the dawn of superintelligent AI would not merely represent another technological upheaval. In their risk calculus, the world should approach it with something like the gravity and caution reserved for nuclear weapons. As Krueger put it, “I do think, on the current trajectory, [the end of humanity] is more likely than not.”

Despite my Burkean predisposition against political interventionism, stringent international controls and careful government monitoring of nuclear weapons are a good idea. Unlike weapons of mass destruction, however, more intelligent software does not self-evidently pose an existential threat. I would not deny that a doomsday scenario is possible, but I have yet to be shown how that would happen. Some say that the time for details is for later, but when basic freedoms are at stake, the burden of proof falls on the interventionist—in my view—to present a case that appeals to both logic and reason.

Climate advocates managed to persuade much of the public to take seriously harms stemming from fossil fuels, even though they are not intuitively dangerous. Figures such as former Vice President Al Gore succeeded, at least in part, by appealing not only to people’s fears but also to their reason. If Krueger and others want to grant the state awesome new powers in the name of AI safety, they must persuade enough of us that an apocalyptic outcome is plausible rather than merely imaginable.

It was with this in mind that I approached my conversations with Krueger. Below is a lightly edited and abridged transcript of two interviews conducted in April and June of this year.


Could you begin with your background in AI research and how it led you to found Evitable? Also, what does the name mean?

Evitable—as in “not inevitable.”

I have been thinking about big-picture questions about the nature of the world and intelligence and human society and how we organize ourselves, including what is wrong with the world because I have seen a lot of problems and have been very cognizant that there are lots of issues in the world since I was quite young.

One of the topics there was artificial intelligence. That is something that I was interested in originally, philosophically, when I was in high school. Then, when I got to college, I learned that it was actually not just science fiction but also a thing that people study, so I became very interested in learning more about it.

When I got into college, I was pretty optimistic about technology and building technologies to help address some of the problems I saw in the world. I got less optimistic about that over time. I think there are still lots of roles for technology in making the world better, but this kind of naïve techno-optimism, where it is just like, “More technology is better,” is something that I really started to question.

In particular with AI, I started to see how it might lead to some really dystopian and dehumanizing futures, with the potential for things like mass surveillance, mass manipulation, and mass automation—really taking humans out of the loop in terms of having any serious decision-making role but then also increasingly scripting their lives for them according to the interests of powerful actors like companies and the government.

Or, more broadly, AI leads down a path of doing things that are not really in anybody’s particular interest. The logic of competitive pressure says, “Oh, we should just keep working really hard to produce more and more stuff,” even if a lot of it is not really making people happy, and we might all be happier if we could spend more time with our families.

Then I also heard about, just from spending a bunch of time on the Internet, people from basically outside the field of AI who were worried about artificial intelligence becoming superintelligent, going rogue, taking over the world, and killing everyone—basically the alignment problem and existential risk.

I found that pretty compelling. At the time, I thought—and I still think—that a lot of these arguments are framed too much in technical terms. I think the underlying issue, or the source of a lot of these problems, is more likely to be social than technical.

To make that a little bit more concrete, you have the idea that AIs are going to ruthlessly optimize for something other than human well-being. I think a lot of people think that is going to happen automatically if we do not build the system right. I think that is plausible, but I think there is also the possibility that people are just going to build the systems that way because that is what is going to be competitive. They are going to build something that is optimized for making profit, for instance.

Then I heard about deep learning in 2012. Before that, I was thinking AI was centuries away. So I was worried about this stuff, but I was thinking more about how to solve these underlying problems of collective action and how to organize society better.

Once I heard about deep learning, I realized that this could all happen in my lifetime. We could have superintelligent AI maybe in a couple of decades. I also saw that this field was about to explode, and it seemed like a great opportunity to get involved and also to learn what the experts thought because these concerns about existential risks seemed to be coming from outside the community of experts.

I was like, “Maybe this is all nonsense, and if I meet with some of the actual leading AI experts, they will be able to explain to me why I should not be worried about that.” Instead, I found out that nobody in the field was really thinking about this stuff or had good counterarguments. Ever since then, I have been very concerned with trying to change that and get more people to understand these concerns, start taking them seriously, and work to address them.

That is kind of my background. Then, I became a professor at Cambridge. I initiated this statement on AI risk that was signed by a bunch of the leading figures in AI, the CEOs of the leading companies, the three most cited scientists of all time—actually, I am not sure if he was number three; he might be—and lots of other AI professors, saying that mitigating the risk of extinction from AI should be a global priority, like nuclear war and pandemics.

That was in 2023, so certainly after ChatGPT came out. Then we started to see the mainstream interest in AI, which continues. But I was hoping for more of a change, more of a reaction from governments. I went to work for the U.K. government on the founding team of the first government body dedicated to addressing this stuff, the U.K. AI Safety Institute [now the UK AI Security Institute].

I quit pretty quickly for a number of reasons, but the main ones were that I was not able to continue doing this kind of public outreach, and also that the work they were doing just did not seem adequate to address the risk. It was more technical research and safety testing, which I think is just not really mature. I think it provides false assurances and is predictably going to break down and fail.

I think we have seen some big signs of that recently with the most recent Anthropic release, where they said, “We kind of just went on based on vibes in the end” because they could not really evaluate the models properly. Also, they were pretty sure the models knew they were being tested. So, you know, hope they are not just fooling us, basically.

You have compared AI risk with nuclear proliferation and climate change. Nuclear destruction is easy to picture; AI extinction is not. What would the path from powerful AI to catastrophe actually look like?

The first thing is: We do not actually know. It is really hard to predict, and, to some extent, it is hard to imagine. It could be something that is very different from anything we have experienced. So that is important to note.

But I think the key thing here probably is that you need to imagine a world where AI is everywhere, and it is increasingly physically embodied and in direct control over physical systems. So it really has more and more of all the power. That is not necessarily one big AI system. It can be a bunch of different AI systems. But the point is that it is everywhere, and it is able to do all of the things that humans can do right now, and more.

It can launch nuclear weapons. It can build and deploy new military technologies. It can decide who gets what resources, what things are built, where those things are built, and who gets to keep them. We rely on the whole industrial society that we have built in order to survive, and AI could be controlling all of that.

If you can see how any existing technology could be a threat to humans, whether that is nuclear weapons, conventional weapons, or bioweapons, AI could be controlling that as well.

Beyond that, AI is going to be—I do not love this crazy analogy—but there is the “country of geniuses in a data center” idea. We are going to have AI systems that are much smarter than people, think much faster, and can develop new technologies at an unprecedented rate. They are going to, again, be in charge of physical systems, so they can design and engineer something, build it, control it, and deploy it.

The details of that are still a bit hard to predict. But I am curious if this gets at it for you—if, when you imagine such a world, it becomes easier to see how we go extinct and how it goes off the rails.

What outcomes concern you most, and where do you think the risk comes from?

The worst-case scenarios are apocalyptic and dystopian and include things like the end of humanity. That is the one that I probably think most about as the thing that I am most concerned about. And I do think, on the current trajectory, that is more likely than not. It is not inevitable, of course.

But there are some other really bad outcomes that are worth mentioning as well, which are situations where we are not literally extinct, but we are disempowered—in some sense, on the periphery of society. Maybe our population is dramatically reduced. Maybe we are confined to Earth, while AI is exploring the rest of the universe. Maybe AI is running society in a way that is treating humans similarly to how we treat other animals, whether that is pets that, in some sense, have a relatively good life or farm animals that may have a terrible life.

Those are the absolute worst-case-type things. Then, there are also situations, which I think are probably going to be transient, but, in the more immediate future, we might have a situation with AI where we do not lose control of it. It might be that our society adapts in such a way that some humans still have power, but maybe it is a very small number of people who have all the power. Maybe nobody is able to get a job anymore, and it is only the people who own AI companies who have any power. Everyone else is sort of left out. They may or may not be provided with the resources they need to survive.

In terms of where the risk is coming from, I would point to two factors: technological and social factors. They are both really important, I think. They play different roles in terms of different bad outcomes.

A lot of people in my circles—let’s say the AI safety community—think almost exclusively about the technical side of things and are almost exclusively concerned with what I would call rogue-AI scenarios. We lose control of an AI system, or AI systems, that act against our interests and attempt to seize power, disempowering us and quite likely leading to our extinction, whether that is directly by killing us or indirectly by rendering the Earth uninhabitable as they build out, most likely, massive computing infrastructure that may do things like directly increase the surface temperature of the Earth and boil the oceans.

Technically, there are major unsolved research problems about how to control AI as it gets more and more intelligent: how to align it or instill a particular motivation in it, how to test it to see if it is safe, and how to understand the processing that it is doing and how it makes decisions and produces outputs.

So there are three or four problems there, basically: alignment, safety testing, and interpretability. These are all big outstanding research problems that people, including myself, have been working on for over a decade. They have received a ton of research attention, and we do not have a plan for how we are going to solve them that we could execute in any known amount of time, given even an infinite budget. We just do not know how hard they are. They could well take decades to solve. They could well be effectively unsolvable for the kind of systems that we currently develop.

I have trouble believing in a literal global takeover scenario. Have you ever been to a developing country? Let’s say you go to a small village in Africa that has not been penetrated by basic technological infrastructure. How is AI going to be in charge?

Right now, humans are a major bottleneck in producing and distributing more machines and more technology. We rely so much on human labor at every step of that process, from mining to manufacturing to transportation. All of that is going to be automated. Once it is automated, and once we have the ability to use that kind of industrial production system to build more AI and build more robots, this can all happen way faster than it has happened in the past.

Then we are also going to be doing tons of research into how to do all that more effectively. So there are a lot of things that we do right now that are very far from what might be the optimal use of resources. We are just going to be removing this huge, huge bottleneck that humans have posed.

I would say humans are very good at being bottlenecks against technological adoption, even if it is against their own best interest. I am just speaking about the one billion people who live in abject poverty, which seems to stem more from incompetence and corruption or willful isolation, like North Korea. Humans seem to be very capable of resisting technology. So I guess that makes me question the notion that AI is going to infect every corner of humanity and potentially kill it.

Look, we are capable of doing that. But there are a few things going on here. One is that other countries do not interfere with these poorer countries as much as they might, partially out of respect for principles of sovereignty, though also they just have a lot of other stuff going on, and they are not run that confidently themselves.

But also, whoever does not use AI becomes less powerful and becomes an easy target for whoever is using AI to exploit or attack or manipulate. It is true that lots of people will not adopt AI, including in rich countries. They will not adopt it aggressively. Those people will be disempowered. And the people who adopt it more aggressively will be empowered, but, by the very nature of this, they will not actually have that power because the whole reason they are getting it is that they are handing it off to AI.

So humans become a conduit for AI decision-making, and the ones who take themselves out of the loop more aggressively end up with more, let’s say, money in their bank accounts. But who is deciding how that money is spent? It is their AI. It is not them. That is why they got so much money in the first place.

There are already a lot of countries and a lot of places that have an overwhelming technological and economic disadvantage, but their sovereignty is still not being violated.

I would say their sovereignty is being violated to some extent. I think that is just the nature of the system. Many regimes are de facto puppet states.

In your conceptualization, do you imagine a proliferation of embodied AI—robots, in other words—playing a major role?

I think that needs to be discussed explicitly because I think you can include it or not include it, and I just want to make sure that people are on the same page about that. I do believe that robotics will be solved to a huge extent. So far, it is lagging AI, and it probably will lag AI, but not by that much. I think we will have superhuman AI probably within four years, and then I think we will get superhuman robotics shortly thereafter—maybe within another four years or less.

I imagine the scenario of robots making robots. That to me makes a lot of sense from a doomer point of view.

Yes, that is exactly what you need to imagine if you want to get very concrete and literal about it. It might not look like that, but that is the sort of thing that it could look like.

Then you have not just robots making robots, but robots designing robots. Robots are doing all of the business side to sustain that and make it very profitable. They own the supply chain—or at least have it distributed across different companies because of antitrust or something—but they close the entire loop, so that the whole thing—from digging stuff out of the ground, refining it, designing a thing, building it, and iterating that whole loop—is happening with no human involvement.

So if the humans just stepped back and did nothing, the robots could literally take over the entire planet and just convert all of the resources they can get their hands on into more robots.

Why could humans not simply pull the plug once that loop became dangerous?

You also have robots running the government and the security forces. And maybe it is not robots building themselves. Maybe there are five or fifty companies involved in this loop, and the loop is still there, but we are not even seeing it. And maybe it is happening in a way that involves a fair number of economic incentives playing a big role.

It is like with climate change. We could pull the plug, so to speak. We could stop burning fossil fuels. But how would we do that? It would require very coordinated action. Similarly, to the extent that AI is very economically valuable, everybody in the system has the incentive to use more AI, give it more power, and give it more autonomy. Then, it is going to be difficult for there to be a clear point of intervention and global coordination.

What does the social side of the risk look like?

I think this is kind of the fundamental dynamic that I was describing: adopt or die, essentially. The problem there is that it does lead to human disempowerment because adoption here means letting AI do all the decision-making.

Every time we come to a point where a human makes a business decision, it is just going to disadvantage that business. If it is a publicly traded company, the shareholders could sue for using human decision-making anywhere. It would be like, “What the hell are you doing? AI knows how to run companies better than humans.”

Then you would be like, “Okay, what is this company doing?” Well, I guess it is maximizing shareholder value. What that actually means, if you take it to the extreme, is pretty unclear. If we have an economy where AI is making all the decisions and human decision-making is considered uncompetitive, it is pretty hard to predict where that ends up.

One way that I think about this, as we talked about in our paper on gradual disempowerment, is that you have all these social systems that are not aligned with human interests. They are somewhat aligned because there are humans who are necessary to run them. Because there is human decision-making involved in running a company or running a government, that provides some alignment. Then they are also kept in check by each other in a balance-of-power arrangement. All of this, put together, sort of works.

But it seems like the human decision-making part—that humans are necessary to run this stuff—is pretty load-bearing. If you just have a bunch of different companies that are all competing for profit or for charitable value or whatever, does that actually continue to reflect human interests? You might think "Yes," because the shareholders are humans. But, again, if the ones who are making the most money and have the most economic power are the ones who are not making their own decisions but are letting AI make decisions for them, then it seems like AI is the one that is really in charge here.

Then it all comes down to questions of alignment: Are these systems doing things that are in the interests of the humans who ostensibly could be, or human systems which should be, in control? But also, how does that all fit together? If nobody is trying to do what is in the collective good, and everyone is optimizing for something like making a successful company, maybe you end up with a bunch of successful companies but no humans alive anymore. They are successful because there are AIs producing and consuming whatever they are building.

This can also go hand in hand with things like the adoption of cyborg technology and uploading and stuff like that. We may also be developing technology that allows people to merge with the machines in some sense. Again, there will be a strong incentive to adopt that technology aggressively. Then, if you essentially let the AI make all the decisions, maybe it starts out with keeping your brain, and then a few years later, you are like, “Why do we not just replace the whole brain?”

Five years ago, I would have doubted that people would willingly let AI make important decisions. But I already see myself and others outsourcing thought—even personal choices—to ChatGPT. That makes this scenario more convincing.

And we are just getting started. The competitive pressure will become much more intense as AI becomes more capable, more autonomous, and better able to act in the real world.

AI agents are already used for personal-assistance and software-engineering tasks. The companies’ game plan is for AI to do anything you can do on a computer. Eventually, you may think, "Nobody is making me automate this, but if I do not, I will be left behind. I may be unemployable." It will be like trying to live as a caveman.

Previous technological revolutions displaced workers and created inequality, but they also produced extraordinary prosperity. I imagine that democratic processes—and possibly policies such as universal basic income—can distribute AI’s gains while preserving human control. So does this necessarily have to end badly?

Not necessarily. What I said is that I think it is more likely than not if we do not course-correct. There are lots of things that we can do to improve things on the margin without stopping the development and deployment of more powerful AI technologies. But I think if we want to actually reduce the risk to an acceptable level, that is clearly what we should be doing.

Right now, the leaders of major world powers should be talking around the clock about how we do this. How do we get to a state where we are sure that nobody is trying to build superintelligence and where we are able to trust each other that we are not trying to do it in secret?

Probably the simplest thing to do is just try to get rid of all of the giant AI data centers, all of the recent chips, and the advanced manufacturing techniques. These things are extremely concentrated and could be paused or eliminated through government actions.

That would leave consumer devices and consumer electronics. That stuff has not improved that much. So it would not be like, “Oh my God, we cannot use computers anymore.” It would just be like: No more AI computers.

That is what I think should happen. But even short of that, if we just keep going on the present course, there is still a good chance that things turn out okay or even great. I am not saying that that is impossible by any means. I just think: Why would we gamble with our future in this way? That is completely reckless and insane. That is more of my perspective on it.

If technical alignment is so difficult, could governments instead require a human to remain in the loop—for example, by prohibiting AI systems from controlling money or making certain decisions?

Basically, the kind of policy you are proposing, I would say, is to keep a human in the loop and not let people outsource their decision-making to AI. You could try to enforce something like that. It would require unprecedented levels of surveillance because, at the end of the day, what is AI? It is software that somebody can run on their own personal computer.

Unless we prevent the types of AI systems that we are worried about from existing or proliferating, there will, based on what we have seen, be an open-source version within roughly a year—or perhaps sooner. 

I think there are really three clear scenarios here. One is where you do not do something like that, and you just kind of let it rip and let people use AI as much as they want. Then, they outsource all their decision-making to AI, and we have to hope that it is sufficiently aligned and also that all of the competing incentives balance out.

We also have to hope that we do not see a race to the bottom of adopting AI that you do not trust, that you may even know is working against your long-term interests, out of a sense that it is necessary in order to survive in the short term because of the competitive pressures, or that it is necessary to at least have a good job or continue to have any standing in society. That is kind of the let-it-rip scenario.

Then, in the AI discourse, I would call the other one “lock it down.” What you talked about, I think, in practice amounts to the lock-it-down thing, where government has to watch everything that everyone does with any computer anywhere. That is pretty dangerous from the concentration-of-power point of view. It is also untested and likely to fail. We have seen attempts at this kind of authoritarianism in the past, and they have all been imperfect.

The third option is what I suggested, which is to try to get rid of advanced AI entirely and get rid of the means for producing advanced AI and then use this time to figure out all of the technical and social questions to our satisfaction instead of just having to slap some stuff together and hope that it works.