This website uses cookies

Read our Privacy policy and Terms of use for more information.

Anthropic Safety Researcher Says There's a Greater Than 10% Chance AI Could Kill All Humans

A senior safety researcher at Anthropic said Tuesday that he believes there is a greater than 10% chance artificial intelligence "could kill all humans" within the next decade, a public statement that came hours after a former Anthropic and OpenAI researcher announced his resignation, citing concerns that leading AI labs are "gambling with our lives," according to CBS News's reporting on the exchange.

What Anthropic's Own Alignment Lead Actually Said

Evan Hubinger, Anthropic's San Francisco-based Alignment Science Lead, posted the warning directly on X in response to his former colleague's resignation announcement. "We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," Hubinger wrote, according to CBS News's reporting. "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." Hubinger separately clarified, according to MSN's reporting on the same post, that he considers the risk from AI models currently in existence to be "low," but that he remains genuinely worried about the trajectory as the technology continues improving toward the ability to self-improve.

The Resignation That Prompted the Response

The exchange began with Jacob Coxon, who spent three years doing pretraining research at both OpenAI and Anthropic, announcing his resignation in a lengthy thread on X. "Neither company is acting responsibly," Coxon wrote, according to Fox Business's reporting on the thread. "They are racing straight to self-improving superintelligence and gambling with our lives." Coxon drew a specific distinction between the two labs' internal cultures: "At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first, they believe no one else will act responsibly, so they must do it themselves, despite the risk."

Key Details Behind the Exchange

Detail

Information

Hubinger's stated risk estimate

>10% chance of AI killing all humans within a decade

Timeframe cited

Next decade

Coxon's tenure

3 years, pretraining research at both OpenAI and Anthropic

Reach of Hubinger's post

9.6 million views

Model withheld from external security testing

Claude Mythos 5.1

UK AISI's response

Confirmed continued collaboration, cited testing OpenAI's GPT-6 Astra "only last week"

A Genuinely Notable Detail About Model Testing Transparency

Coxon's resignation thread surfaced a specific, concrete disclosure worth understanding. In a corporate blog post the week prior, Anthropic revealed the company had not shared its latest AI model, Claude Mythos 5.1, with security bodies outside the United States, including the U.K.'s AI Security Institute, widely regarded as a world-leading body for testing frontier AI model risks, according to CBS News's reporting. A Cabinet Office spokesperson responded on the UK government's behalf: "The AI Security Institute continues to collaborate closely with industry partners, including Anthropic, to make models safer," noting the institute had tested OpenAI's GPT-6 Astra "only last week" before its public release. The spokesperson added: "These risks do not stop at national borders and no country can tackle them alone."

Anthropic's Own Response to the Broader Concern

Anthropic itself addressed the underlying issue directly in a separate blog post, according to CNBC's reporting on the exchange: "If systems are capable of fully building their own successors, the ways we secure them, monitor them, and shape their behavior all grow much more important." That framing reflects an underlying tension in Anthropic's own public positioning worth understanding: the company continues to raise significant capital and pursue an expected public listing even as its own senior safety researchers voice these kinds of concerns publicly and unprompted.

Why This Isn't a Genuinely New Concern, But the Public Framing Is Notable

The notion that frontier AI models could pose an existential risk to humanity isn't new, and executives at both OpenAI and Anthropic have made similar statements in the past. What distinguishes this moment is the directness and specificity of the exchange, a resignation citing genuine moral concern, followed by a current senior employee publicly agreeing with a specific numerical risk estimate rather than offering vague reassurance. This connects directly to the broader pattern of AI safety incidents and concerns we've tracked closely this month, including experts calling the Hugging Face hack a "warning shot" and OpenAI reportedly sitting on a second undisclosed rogue agent incident.

Why This Matters for Business

This exchange is worth understanding for any business making long-term AI vendor and infrastructure decisions, since it represents current, senior safety personnel at one of the industry's most safety-focused labs publicly expressing genuine uncertainty about whether alignment challenges for increasingly capable AI systems are actually being solved, not merely theoretical outsider criticism.

For businesses evaluating AI governance and risk management internally, this public disagreement between AI insiders about the pace of development versus the maturity of safety solutions is worth factoring into any long-term technology strategy that assumes continued, uninterrupted AI capability growth without corresponding safety breakthroughs.

Frequently Asked Questions

What did the Anthropic researcher actually say about AI risk?
Evan Hubinger, Anthropic's Alignment Science Lead, said he believes there is a greater than 10% chance AI could kill all humans within the next decade, while clarifying he considers the risk from currently existing models to be low.

Why did Jacob Coxon resign from Anthropic?
Coxon, who previously worked at both OpenAI and Anthropic, resigned citing concerns that both companies are racing toward self-improving superintelligence without having solved fundamental safety and alignment challenges.

Did Anthropic share its newest model with international safety regulators?
Not entirely. Anthropic disclosed it had not shared its latest model, Claude Mythos 5.1, with security bodies outside the United States, including the UK's AI Security Institute, though the UK government said it continues collaborating with Anthropic on safety testing more broadly.

The Fast Version

Anthropic Alignment Science Lead Evan Hubinger said publicly that he believes there is a greater than 10% chance AI could kill all humans within the next decade, responding to former Anthropic and OpenAI researcher Jacob Coxon's resignation, in which Coxon accused both companies of racing toward self-improving superintelligence without adequate safeguards. The exchange surfaced a separate disclosure that Anthropic had not shared its newest model, Claude Mythos 5.1, with international safety regulators including the UK's AI Security Institute, though the UK government said its collaboration with Anthropic continues. The public exchange adds to mounting concern from within the AI industry itself about whether alignment and safety research is keeping pace with rapidly advancing model capability.