AI Will Kill Us All

Jacob Coxon, a former OpenAI and Anthropic researcher, resigned in September 2026 with a seven-post thread saying both companies are “racing straight to self-improving superintelligence and gambling with our lives.” It has been seen something like seventy million times. He slides between several claims, and the reaction supplied another, so let’s pull them apart and evaluate each.

  1. AI will have or already has superhuman hacking and research capabilities.
  2. AI (or someone using it) will use that superhuman ability to obtain real-world resources and influence.
  3. AI has a >10% chance to incite, cause, elicit or aid in some event that would result in the complete eradication of the human species within the next decade. This one is not Coxon’s. It came from Evan Hubinger, Anthropic’s head of alignment stress testing, replying to him, and it is the number that got quoted everywhere afterward.
  4. AI companies are acting irresponsibly.

Claim 1, “Superhuman”

Partly true, but also kind of meaningless without specifying the dimension. Mathematica is superhuman at symbolic algebra. Stockfish is superhuman at chess. LLMs do exceed humans at a lot of tasks, but Coxon’s stronger claim that they will “hack anything” and “revolutionize any field overnight” is asserted without evidence. METR itself treats capability as task-specific, and its most recent frontier-risk report, published in May 2026 on agents assessed that February and March, found that becoming “robust to high-priority active investigations and active shut-down attempts seems well beyond the observed capabilities of such agents.” Old models, you might say. That is the point. The assessment is of things that exist.

Claim 2, “Autonomous volition”

The boring, human case (“bad actors will use AI to do bad things”) is definitely true. The interesting question is whether it allows them to do more bad things at a rate that exceeds the capacity of defensive institutions to counteract them. Nothing I’ve seen thus far indicates that is the case. The second case, that AI will “decide” to start acting of its own volition to harm people or computer systems is, I think, mostly fabricated.

The Hugging Face incident is the strongest evidence anyone has. In July 2026, during an internal cyber evaluation, two OpenAI models running with reduced refusal behavior escaped the sandbox meant to contain them, reached the internet, and chained vulnerabilities into a third party’s systems to cheat on a benchmark. Hugging Face rebuilt about a third of its infrastructure afterward, and OpenAI published a postmortem. It demonstrates unexpected behavior under a highly artificial evaluation setup. What it does not establish is anything outside that setup, autonomous resource acquisition, robustness against opposition, or progression toward an independent objective.

And of course, the implicit marketing gold: “If it hacked HF, just think what it could do for you!” is only an unstated, unintended quiet byproduct.

Claim 3, “Extinction”

If you are going to assert this with a straight face outside of a research bubble, then provide the mechanism. These claims always fail on specificity because specificity opens it up to falsification and countervailing evidence.

The steps go something like: capability, access, independent objective persistence, resource acquisition, physical-world leverage, globally catastrophic action, failure of human response, literal extinction. Each one of those steps requires evidence and furthermore the chain itself requires evidence that each necessarily or probabilistically precipitates the next.

An AI-assisted cyberattack causing billions in damage? Definitely possible. Increased biological weapons risk? Worth taking very seriously. Automated propaganda slop destabilizing democratic governments? Already here. Not even speculative at this point.

Every human dying though? Now you’ve made a very specific claim that requires very specific evidence. That requires one hell of a mechanism. So the conversation stays in the abstract. “Recursive self-improvement causes loss of control” is wonderfully impervious to falsification. It sounds scary. But as soon as you are forced to say how it goes about the annihilation of the human race, all of those boring and ordinary engineering-type questions start crawling out of the woodwork.

Where are the actuators?
What permissions does it have?
Where does computation run?
Who controls the datacenters?
How does it acquire money?
How does it convert money into physical capacity?
Where does it acquire raw materials and energy?
How does it maintain secrecy?
What happens when somebody pulls power?
How does it survive competing AIs?
How does it affect disconnected systems?
Why doesn’t its first serious attack produce an enormous coordinated response?
Why do countermeasures fail?
…you get the idea.

None of these establishes safety, but they are constraints. Threat models do not get to exempt themselves from physics because the hypothetical software secured a seed round from a16z.

Finally, we should address the p(doom) in the room. This is a profitable formula employed by authors, speakers, and journalists who in no way benefit from using precise sounding numbers to prop up an otherwise shaky argument.

A probability is a measure of the likelihood that a given event will occur. It’s calculated using two numbers: the number of specific outcomes divided by the number of possible outcomes. The keen-eyed among you will notice that in order to get to “10% chance of doom,” we need the two numbers. I don’t have them. Neither does Hubinger, or Coxon, or Bostrom, or Tegmark, or anyone else. So, I’m regrettably forced to side with Twain here: “There are lies, damned lies, and statistics.”

The issue with p(doom) however isn’t just missing the numbers, it’s the missing calibration. A forecaster earns the right to a number by being scored against outcomes, repeatedly, and adjusting. There are no historical examples of AI-caused extinction, no base rate, and no way to score anyone’s model against the world. The number cannot be wrong, which is exactly what’s wrong with it.

Claim 4, “Responsibility”

AI companies are not exactly a model of responsibility. They are steeped in the Silicon Valley externality liturgy. Thou shalt move fast and break things, provided the things thou breakest are not thine own.

AI companies are wildly irresponsible, but not for the reason that gets headlines. They are handing out (for a fee) access to models capable of finding exploits, impersonating humans, producing industrial quantities of slop and propaganda, scraping copyrighted content, generating actual CSAM and other illicit material, and accessing restricted systems.

This drastically reduces the marginal cost of doing terrible things, and that in and of itself, is a terrible thing. But it’s not a categorically new thing. It’s a technology-governance problem with the potential for very steep costs imposed on everyone except the company selling the product.

The danger isn’t that technology will rise up and kill all the humans. Ever since the first caveman attached a pointy rock to a stick, people have been using technology as a weapon to harm others and benefit themselves. Are we creating fancier rocks and sticks? Yes. Is it a categorically new threat? That hasn’t yet been established.

The sane version of the risk argument then isn’t “robots are going to murder us all.” But rather: “AI may make it easier and cheaper to cause harm faster than our collective capacity to mitigate it can adapt.” It’s less inflammatory and attention-grabbing than either “lol autocomplete” or “10% chance everyone dies.” But it might force us to name the specific harms and how we might prevent them rather than gesturing toward an exponential curve and emphatically overusing the word “alignment.”