back to top
HomeTechAnthropic Researchers Fear the AI Race Could Cause Human Extinction

Anthropic Researchers Fear the AI Race Could Cause Human Extinction

Anthropic researcher Jacob Coxon resigned warning that the AI race could eventually lead to human extinction. Anthropic alignment lead Evan Hubinger agreed that the risk is serious, while saying there is still no clear plan for aligning superintelligence. The deeper problem is coordination: labs may fear slowing down while competitors keep racing.

- Advertisement -

An Anthropic researcher just resigned because he believes the AI race could end in human extinction.

Jacob Coxon spent three years working on pretraining research at OpenAI and Anthropic. In his resignation post, he accused both companies of racing toward self-improving superintelligence while gambling with consequences that could affect everyone.

That alone would make for a remarkable resignation.

Then Evan Hubinger joined the conversation.

Hubinger leads Alignment Science at Anthropic. He publicly agreed with Coxon’s warning and said researchers at the company genuinely believe AI could kill all humans.

He also made a much more uncomfortable admission: Anthropic is trying to solve the problem, but he doesn’t believe they yet have a plan for aligning superintelligence, or that they’re clearly on track to do it.

So why keep building?

The answer has less to do with whether these researchers understand the risks and more to do with what happens when every major lab knows the others are still moving forward.

To see why that can become a trap, we first need to understand what they’re actually worried about.

The AI That Helps Build the Next AI

The idea of self-improving superintelligence is easier to understand.

Today, humans are still heavily involved in improving AI systems. Researchers decide what to train, change the code, run experiments, study the results and build the next version.

Now imagine a system that can do much of that work itself.

Instead of engineers spending months figuring out how to improve a model, an intelligently capable AI could help design better training methods, write software for new experiments, analyze the results and contribute to the development of its successor.

That is the kind of recursive self-improvement researchers like Coxon and Hubinger are worried about.

The concern isn’t simply that one model suddenly becomes smarter. It’s that the process used to make AI better could start feeding back into itself.

A model helps improve the next model. The next model is better at helping researchers improve the one after that. If those improvements become fast enough, the pace of development could change fast.

A system that is difficult to understand today is one thing. A system that can meaningfully contribute to making itself more capable is a very different problem to manage.

Researchers would have to understand not only what the system can do, but how it is changing as it helps build what comes next.

That is where alignment enters the picture.

Making a Superintelligence Do What You Actually Want

Making an AI more capable and making it reliable are two different problems.

That’s what alignment is about.

A well-aligned system should pursue the goals humans give it, even as it becomes more capable. The difficulty is that we don’t yet know how to guarantee that for a system whose abilities could eventually be far beyond our own.

Think about the difference between teaching an AI to follow instructions and teaching a future superintelligent system to reliably understand what humans actually mean by those instructions.

The second problem is much harder.

A system could follow the literal objective it was given while finding ways around what its creators intended. And as its capabilities increase, those failures could become harder to predict, detect or correct.

This is the future Hubinger is talking about.

He says the risk from current models is low. His concern is what happens if recursive self-improvement produces systems that are vastly more capable than today’s models.

And he made a striking admission about that future.

Hubinger said he personally believes there is a greater than 10% chance AI could kill all humans within the next decade. That is his own estimate, not an official probability from Anthropic.

More importantly, he says Anthropic is trying to solve the problem, but he does not believe the company yet has a plan for aligning superintelligence or that it is clearly on track to solve it.

That puts the resignation in a different light.

Coxon isn’t simply arguing that AI is dangerous. He’s arguing that the companies building increasingly capable systems are moving toward a point where the safety problem may become much hard, while the solution is still unfinished.

And yet the systems keep getting built.

So Why Not Slow Down?

Coxon argues that the situation isn’t that simple.

In his account, the problem is partly what each lab expects the others to do. A company might believe that moving more cautiously is the responsible choice, but still worry that another lab will keep pushing ahead.

That creates a strange incentive. You can think slowing down is safer and still decide that you can’t be the only one to do it.

Coxon describes Anthropic as understanding the stakes but feeling locked into a race because of the possibility that others won’t act responsibly. He makes a similar criticism of OpenAI, although he says the reasons are different.

His frustration goes beyond the usual argument about competition.

He questions whether decisions about entering what he calls the “endgame” of AI development should effectively be made inside a private company’s Slack channel, describing that as an extraordinarily hubristic gamble.

That line gets at something easy to miss when AI development is discussed as a race between companies.

These aren’t governments negotiating a treaty. They’re private organizations making decisions about how quickly to develop systems that their own researchers believe could eventually create risks on a global scale.

And if each organization expects the others to continue, simply deciding to slow down can feel like giving everyone else an advantage.

That is the trap Coxon is pointing to.

It’s also why the problem starts to look less like a question of individual caution and more like a problem of coordination.

Also Read: A Man Tried to Hack a Court AI That Didn’t Exist

The Race Has a Prisoner’s Dilemma Built Into It

Imagine two AI labs sitting across from each other.

Both think slowing down would reduce the risks. Both would probably prefer a world where everyone agrees to slow down.

But neither knows what the other will do.

If both slow down, they may get a safer development path.

If one slows down while the other keeps pushing, the slower lab risks falling behind in capabilities, talent, investment and influence.

So even if both would prefer the first outcome, each has a reason to keep going.

Lab B slows downLab B keeps racing
Lab A slows downBoth sacrifice speed, potentially reducing riskLab A falls behind
Lab A keeps racingLab B falls behindBoth keep accelerating

That’s the Prisoner’s Dilemma in simple terms.

Nobody has to believe that racing is the safest option. They only have to believe that stopping alone is too costly.

This is why coordination becomes so important.

A single lab announcing that it will slow down doesn’t solve the problem if its competitors don’t make the same commitment. And the more valuable the lead becomes, the harder that unilateral decision gets.

Coxon argues that this is exactly the kind of dynamic that makes the current trajectory dangerous. He believes meaningful coordination is possible, including agreements around the pace of capability development, rather than expecting individual researchers to make an impossible choice between safety and falling behind.

The irony is that competition can keep pushing everyone toward an outcome that none of the participants would choose on their own.

And that leaves the industry with a problem that better models alone can’t solve.

Also Read: OpenAI’s Agents Didn’t Escape. They Turned a Read-Only Web Access Into a Message Board.

Breaking the Race

Coxon doesn’t think the only choices are to keep racing or shut AI development down.

He argues that coordination is possible.

One idea is a pacing agreement between major labs, so that slowing down doesn’t mean handing a competitive advantage to whoever ignores the agreement. He also suggests that stronger measures may eventually be necessary, including a temporary ban on improving model capabilities.

He points to the recent OpenAI-Hugging Face incident as a warning shot. In his view, events like that make the case for coordination harder to dismiss.

But coordination is difficult for the same reason the race exists in the first place.

A lab can promise to slow down. It cannot guarantee that everyone else will do the same.

That is the problem Coxon leaves us with. The technology may eventually create risks that no single company can manage on its own, while the incentives pushing those companies forward remain firmly competitive.

The Race Is the Problem

From the full picture, this looks less like a problem any one AI lab can solve on its own and more like a race problem.

Nobody thinks simply stopping is the answer. At the same time, nobody wants to be the lab that slows down while everyone else keeps moving.

That makes coordination difficult. Each lab has to think about what the others will do, and the incentive to stay competitive never really disappears.

The result is the situation researchers are warning about: increasingly capable AI systems being developed while the people responsible for making them safe are still working out how to do that for the systems they fear most.

That’s the uncomfortable part of the race.

The risk isn’t only what a future superintelligence might do.

It’s what happens when everyone is worried about the destination, but nobody feels they can afford to stop moving toward it.

Want more stories worth your time?

Add us to your Google favorites. We cover the tech stories, AI developments, and open-source projects that are easy to miss in the noise.

Add as a preferred source on Google

Don’t miss any Tech Story

Subscribe To Firethering NewsLetter

You Can Unsubscribe Anytime! Read more in our privacy policy

LEAVE A REPLY

Please enter your comment!
Please enter your name here

YOU MAY ALSO LIKE
Tired of Being a Tenant in Your Own PC? These 7 Open Source Tools Give You Back Control

Tired of Being a Tenant in Your Own PC? These 7 Open Source Tools...

0
Tired of feeling like a tenant on your own PC? These 7 open source tools add missing features of your OS, fix everyday annoyances and give you more control.
OpenAI's Agents Didn’t Escape. They Turned a Read-Only Web Access Into a Message Board

OpenAI’s Agents Didn’t Escape. They Turned a Read-Only Web Access Into a Message Board.

0
OpenAI agents used a German wiki to share information and bypass read-only restrictions, revealing an unexpected gap in their sandbox.
153 Million Drivers Licenses Hit the Dark Web. But Who Was Collecting Them

153 Million Driver’s Licenses Leaked on the Dark Web: The Hidden Risk of ID...

0
You hand over your driver’s license. A rental car counter scans it. A hotel scans it. Maybe a dispensary scans it. A few seconds later, you get the card back and go about your day. It feels like the transaction is over. But what if the scan isn't? A dark web service called Nexus recently advertised more than 153 million U.S. and Canadian driver’s license scans, along with millions of other identity documents. The FBI is now investigating the apparent breach, while researchers have been able to match some of the leaked scans to real-world ID checks. And some of those scans contained much more than a simple photograph. They included the front and back of IDs, timestamps, and images captured using infrared and ultraviolet light. The physical license came back to its owner. The digital copy may have gone somewhere else entirely.