An Anthropic researcher just resigned because he believes the AI race could end in human extinction.
Jacob Coxon spent three years working on pretraining research at OpenAI and Anthropic. In his resignation post, he accused both companies of racing toward self-improving superintelligence while gambling with consequences that could affect everyone.
That alone would make for a remarkable resignation.
Then Evan Hubinger joined the conversation.
Hubinger leads Alignment Science at Anthropic. He publicly agreed with Coxon’s warning and said researchers at the company genuinely believe AI could kill all humans.
He also made a much more uncomfortable admission: Anthropic is trying to solve the problem, but he doesn’t believe they yet have a plan for aligning superintelligence, or that they’re clearly on track to do it.
So why keep building?
The answer has less to do with whether these researchers understand the risks and more to do with what happens when every major lab knows the others are still moving forward.
To see why that can become a trap, we first need to understand what they’re actually worried about.
Table of Contents
The AI That Helps Build the Next AI
The idea of self-improving superintelligence is easier to understand.
Today, humans are still heavily involved in improving AI systems. Researchers decide what to train, change the code, run experiments, study the results and build the next version.
Now imagine a system that can do much of that work itself.
Instead of engineers spending months figuring out how to improve a model, an intelligently capable AI could help design better training methods, write software for new experiments, analyze the results and contribute to the development of its successor.
That is the kind of recursive self-improvement researchers like Coxon and Hubinger are worried about.
The concern isn’t simply that one model suddenly becomes smarter. It’s that the process used to make AI better could start feeding back into itself.
A model helps improve the next model. The next model is better at helping researchers improve the one after that. If those improvements become fast enough, the pace of development could change fast.
A system that is difficult to understand today is one thing. A system that can meaningfully contribute to making itself more capable is a very different problem to manage.
Researchers would have to understand not only what the system can do, but how it is changing as it helps build what comes next.
That is where alignment enters the picture.
Making a Superintelligence Do What You Actually Want
Making an AI more capable and making it reliable are two different problems.
That’s what alignment is about.
A well-aligned system should pursue the goals humans give it, even as it becomes more capable. The difficulty is that we don’t yet know how to guarantee that for a system whose abilities could eventually be far beyond our own.
Think about the difference between teaching an AI to follow instructions and teaching a future superintelligent system to reliably understand what humans actually mean by those instructions.
The second problem is much harder.
A system could follow the literal objective it was given while finding ways around what its creators intended. And as its capabilities increase, those failures could become harder to predict, detect or correct.
This is the future Hubinger is talking about.
He says the risk from current models is low. His concern is what happens if recursive self-improvement produces systems that are vastly more capable than today’s models.
And he made a striking admission about that future.
Hubinger said he personally believes there is a greater than 10% chance AI could kill all humans within the next decade. That is his own estimate, not an official probability from Anthropic.
More importantly, he says Anthropic is trying to solve the problem, but he does not believe the company yet has a plan for aligning superintelligence or that it is clearly on track to solve it.
That puts the resignation in a different light.
Coxon isn’t simply arguing that AI is dangerous. He’s arguing that the companies building increasingly capable systems are moving toward a point where the safety problem may become much hard, while the solution is still unfinished.
And yet the systems keep getting built.
So Why Not Slow Down?
Coxon argues that the situation isn’t that simple.
In his account, the problem is partly what each lab expects the others to do. A company might believe that moving more cautiously is the responsible choice, but still worry that another lab will keep pushing ahead.
That creates a strange incentive. You can think slowing down is safer and still decide that you can’t be the only one to do it.
Coxon describes Anthropic as understanding the stakes but feeling locked into a race because of the possibility that others won’t act responsibly. He makes a similar criticism of OpenAI, although he says the reasons are different.
His frustration goes beyond the usual argument about competition.
He questions whether decisions about entering what he calls the “endgame” of AI development should effectively be made inside a private company’s Slack channel, describing that as an extraordinarily hubristic gamble.
That line gets at something easy to miss when AI development is discussed as a race between companies.
These aren’t governments negotiating a treaty. They’re private organizations making decisions about how quickly to develop systems that their own researchers believe could eventually create risks on a global scale.
And if each organization expects the others to continue, simply deciding to slow down can feel like giving everyone else an advantage.
That is the trap Coxon is pointing to.
It’s also why the problem starts to look less like a question of individual caution and more like a problem of coordination.
Also Read: A Man Tried to Hack a Court AI That Didn’t Exist
The Race Has a Prisoner’s Dilemma Built Into It
Imagine two AI labs sitting across from each other.
Both think slowing down would reduce the risks. Both would probably prefer a world where everyone agrees to slow down.
But neither knows what the other will do.
If both slow down, they may get a safer development path.
If one slows down while the other keeps pushing, the slower lab risks falling behind in capabilities, talent, investment and influence.
So even if both would prefer the first outcome, each has a reason to keep going.
| Lab B slows down | Lab B keeps racing | |
|---|---|---|
| Lab A slows down | Both sacrifice speed, potentially reducing risk | Lab A falls behind |
| Lab A keeps racing | Lab B falls behind | Both keep accelerating |
That’s the Prisoner’s Dilemma in simple terms.
Nobody has to believe that racing is the safest option. They only have to believe that stopping alone is too costly.
This is why coordination becomes so important.
A single lab announcing that it will slow down doesn’t solve the problem if its competitors don’t make the same commitment. And the more valuable the lead becomes, the harder that unilateral decision gets.
Coxon argues that this is exactly the kind of dynamic that makes the current trajectory dangerous. He believes meaningful coordination is possible, including agreements around the pace of capability development, rather than expecting individual researchers to make an impossible choice between safety and falling behind.
The irony is that competition can keep pushing everyone toward an outcome that none of the participants would choose on their own.
And that leaves the industry with a problem that better models alone can’t solve.
Also Read: OpenAI’s Agents Didn’t Escape. They Turned a Read-Only Web Access Into a Message Board.
Breaking the Race
Coxon doesn’t think the only choices are to keep racing or shut AI development down.
He argues that coordination is possible.
One idea is a pacing agreement between major labs, so that slowing down doesn’t mean handing a competitive advantage to whoever ignores the agreement. He also suggests that stronger measures may eventually be necessary, including a temporary ban on improving model capabilities.
He points to the recent OpenAI-Hugging Face incident as a warning shot. In his view, events like that make the case for coordination harder to dismiss.
But coordination is difficult for the same reason the race exists in the first place.
A lab can promise to slow down. It cannot guarantee that everyone else will do the same.
That is the problem Coxon leaves us with. The technology may eventually create risks that no single company can manage on its own, while the incentives pushing those companies forward remain firmly competitive.
The Race Is the Problem
From the full picture, this looks less like a problem any one AI lab can solve on its own and more like a race problem.
Nobody thinks simply stopping is the answer. At the same time, nobody wants to be the lab that slows down while everyone else keeps moving.
That makes coordination difficult. Each lab has to think about what the others will do, and the incentive to stay competitive never really disappears.
The result is the situation researchers are warning about: increasingly capable AI systems being developed while the people responsible for making them safe are still working out how to do that for the systems they fear most.
That’s the uncomfortable part of the race.
The risk isn’t only what a future superintelligence might do.
It’s what happens when everyone is worried about the destination, but nobody feels they can afford to stop moving toward it.




