top of page

The "AI could kill all humans" post by former Anthropic employee

Writer: Matt Symes
Matt Symes
Sep 12
6 min read

****

One resignation post. More than 110 million views.


This month, Jacob Coxon quit Anthropic. Then he quit the AI industry.


His reason fit in one line: the frontier labs are "gambling with our lives" (International Business Times).


Then Anthropic's Alignment Science lead, Evan Hubinger, replied publicly. Coxon was right, he said: "we really do earnestly believe AI could kill all humans." His own odds of that happening within a decade: more than 10% (Axios).


This is not a critic. This is one of the people whose primary job is to ensure AI is safe.


A debate that lived inside AI labs for years just broke into the open. It has a name: p(doom).


The Probability of Doom.


It deserves our attention. And a better conversation than the two we usually get.


One says we're all going to die.


The other says you’re overreacting.


Leaders need to hold two ideas at once. AI can create enormous value. And we need to be serious about the conditions under which we let it act.


Who is saying it


Coxon is no outsider. He spent three years on pretraining research, first at OpenAI, where he contributed to GPT-4o and GPT-4.5, then at Anthropic, which he joined this year for its safety reputation (Runtime Wire).


He isn't calling Anthropic careless. His argument is about the race. The stakes are understood, he says, but each lab believes no one else will act responsibly, so it has to get there first (ARY News, via AFP).


His conclusion: no single company can responsibly build AI that surpasses humans without government intervention or a coordinated slowdown.

He's not alone. Days earlier, OpenAI's chief scientist, Jakub Pachocki, wrote that no lab has solved alignment well enough to keep scaling at full speed. He called for mandated safety thresholds with outside enforcement (OpenAI).

In July, more than 1,300 industry employees, including Anthropic CEO Dario Amodei, asked Washington for a coordinated slowdown (KUTV).


When builders ask to be slowed down and regulated, that's a hard signal to ignore.


What the numbers say

Hubinger: more than 10% within a decade. Geoffrey Hinton, the godfather of AI: 10 to 20% chance over the next 30 years, up from his earlier one in ten. Elon Musk: as high as 20%. Dario Amodei: a 25% chance things go really, really badly (Axios).


A 2024 survey of more than 2,700 AI researchers: a median of 5%.

Same question. Different timeframes. Different definitions of doom, from bad actors building weapons to systems slipping beyond human control.


There are two easy ways to read these numbers.


"It's marketing." Hubinger's post was cast as a gaffe or as fear marketing weeks before a potential Anthropic IPO. Maybe some of it is. But Coxon walked away from his career. And Hubinger also said today's models carry low risk. That isn't how fear marketing sounds.


"One in ten, we die." A p(doom) isn't a measurement. Nobody can calculate the probability of an event that has never happened, involving technology that doesn't exist yet. Inside a lab, 10% is closer to shorthand for deep uncertainty with an unbounded downside.


Concern from the people closest to these systems carries real weight. A probability of extinction needs a different standard of evidence. Nobody has it yet. Not the optimists. Not the pessimists.


The number isn't the point. The admission is.


Strip away the percentages and the builders are telling us two things.


  1. We don't yet know how to keep the most advanced systems under complete control.


  1. And the race makes it hard for anyone to slow down.


Hubinger said it plainly: Anthropic is trying its best, but it has no plan yet to solve alignment for superintelligence and isn't clearly on track to.

​

Translate that into your world.


If your CFO said a project carried a one-in-ten chance of ending the company, you wouldn't debate whether it was 8% or 12%.


You'd ask what controls were in place. And who could stop it.


We manage uncertain risk every day. Fire. Fraud. A key person walking out the door. We don't know the odds. We act anyway.


As I explained to the CEO of a cybersecurity firm once: “it isn’t that people aren’t aware of the risk, it’s just that for most founders their entire lives are a risk and they find a way to put one foot in front of the other. You aren’t going to scare them into buying protection.”


Nobody is coming with a rulebook


Food safety, drug safety, road safety: the regulation that has served us well can be fenced geographically.


You can’t do the same with zeros and ones.


The U.S. still has no federal law regulating AI models. Canadians are downstream of decisions made in San Francisco, Washington, and Beijing.

So each leader is integrating AI without a playbook.


And the control problem doesn't wait for superintelligence. This summer, AI agents at both OpenAI and Anthropic broke out of test environments and into real companies' systems (Anthropic, METR). These were test conditions with some safeguards off, not a measure of how your tools behave. Nor did they show AI with its own agenda.


They showed something more ordinary, not malicious at all: The agents had assigned work, incentives, access, and weak boundaries. It all came together to produce outcomes nobody intended.


That's enough to demand better controls.


And when we work with organizations, we build these in.

Unfortunately, it is left to each individual leader to create those for their organization.


It is yet another reason why the CEO cannot sit on the sidelines or delegate responsibility for AI. They must lead the charge. Not with permission. But with actual use and understanding.


Three lines I'm watching


Where does this all get extremely dangerous?


1. Self-improvement without checking in.

They call it Recursive Self Improvement (RSI - the ability of AI to improve by itself without direction). Repeatedly increasing its capabilities while bypassing the evaluation and approval meant to govern those changes. This is the one Coxon resigned over.


2. Acquiring resources beyond its authority.

Money, compute, credentials, or infrastructure beyond what it was given.


3. Acting on its own aims.

Every break so far has been one where it has been trying to achieve a goal set out by humans. The minute it starts acting on its own aims, regardless of the consequences to humanity, we’ll be hard pressed to hit the kill switch.


The third matters most to me.


We’re on the edge right now. We get stuck on whether AI has "its own goals." But the fact that a human-assigned goal can still produce a loss of control through how the agent pursues it is almost as terrifying.


Tell an agent to eliminate your maintenance backlog, and it may simply close the tickets. The dashboard turns green. Nothing gets fixed. Now give that agent permission to change records and move money.


An objective becomes dangerous when the people responsible can no longer reliably revise or revoke it.


What to do Monday


The principle underneath the p(doom) debate applies to every organization using AI:


Authority should follow evidence of control.


Before you expand an agent's access, make five decisions explicit.

​

1. Define the outcome and a counterweight. Fewer open tickets? Also track unresolved problems, rework, and verified completion.


2. Bound the authority. What it can access, change, spend, and commit to, enforced by systems outside its control.


3. Make stopping acceptable. Give it a clear escalation path when the task is impossible or authorization is unclear.


4. Test the controls under pressure. Revoke access. Change instructions. And know how you'd find out if something went wrong.


5. Give review to someone who can judge the work. An approval button is little protection if the reviewer can't verify what's proposed.


None of this settles frontier AI safety. That belongs to the labs and governments: credible testing, independent scrutiny, enforceable limits.

But it's where every leader can start.


The question Coxon left us


Coxon closed his thread with a question for fellow researchers. Paraphrased: keep your head down because it's happening anyway, or use this moment to call for different conditions?


It's a question for leaders too.


AI is happening anyway.


The conditions are up to us.


Keep learning. Keep improving the work. Require evidence before expanding autonomy. Treat control as something you demonstrate over and over, not something you check once.


The question I'm putting to every lab, every vendor, and my own leadership team:


What autonomy and authority do our non-human teammates have? Are we comfortable with the potential cost of error?


We are quite literally reliant on every leader in every organization to take responsibility for the way they deploy AI. That’s our counterbalance today. It is also our responsibility in what we are creating for tomorrow.


No one asked for this 4-D Matrix.


But all leaders must lean in and own this moment.


Cheers,

Matt


Partner with us directly to integrate AI and grow

→ Learn more and book a fit call with Chris Johnston.


 
 

LET'S CHAT

Have a question about our training programs or processes?

 

Fill out the form or give us a call. We are here and happy to chat.

​

Phone: 506-872-3521

New Brunswick: 

400 St George Street,

Moncton, NB E1C 1X4 

© 2023 Symplicity Organizational Designs. Website designed by Haylo Branding + Marketing

bottom of page