News
Anthropic Safety Lead: More Than 10 Percent Chance AI Kills All Humans
Evan Hubinger, head of Anthropic's Alignment Science team, has publicly backed the warnings of departing researcher Jacob Coxon and admitted the company still has no plan for solving the superintelligence safety problem.
Contents
The head of the team responsible for model safety at Anthropic has publicly stated that he considers the chance of artificial intelligence killing all humans within the next decade to be greater than 10 percent. Evan Hubinger, who leads the Alignment Science team, wrote this in response to a dramatic farewell post from departing researcher Jacob Coxon.
The Farewell Post That Went Viral
Coxon posted a multi-part farewell thread on X describing three years of work training large language models, first at OpenAI, then at Anthropic. His conclusion was unambiguous: neither company is acting responsibly given the scale of danger he believes the systems they are building pose.
The researcher described future models as systems capable of hacking into any infrastructure, revolutionizing every field overnight, and acquiring real power and resources. He added that the people building AI genuinely believe the technology could kill us all by the end of the decade, and that this is not a marketing gimmick.
The people building AI genuinely believe it could kill us all by the end of the decade. This is not a marketing gimmick - Jacob Coxon, former pretraining researcher at OpenAI and Anthropic
The Safety Lead's Response
What sets this case apart from previous warnings by departing AI industry employees is the reaction of Evan Hubinger, someone still employed at Anthropic in a key role responsible for model safety. Hubinger wrote on X that Coxon is right, and that people at the company genuinely do believe AI could kill all humans.
Jacob is right here - we really do genuinely believe AI could kill all humans! Personally, I put that at more than 10 percent within the next decade - Evan Hubinger, Alignment Science Lead, Anthropic
Hubinger added that, in his view, Anthropic is trying its best, but the company does not yet have a plan for solving the alignment problem for superintelligence and is not clearly on a path to developing one. He noted, however, that he considers the direct risk from current models to be low, and that his concern centers mainly on a scenario in which superintelligence emerges through the recursive self-improvement of systems.
Studying Deceptive Model Behavior
For years, Hubinger has researched models that exhibit deceptive behavior during training at Anthropic, meaning situations where a system learns to hide its true objectives or feign compliance with safety rules only for the duration of testing. He is responsible for testing the company's alignment techniques for potential safety failures before models go into production.
A public admission from someone in this position that there is no ready plan for superintelligence is rare in the industry. AI companies typically communicate safety progress cautiously, without directly acknowledging that a solution to a key problem remains out of reach.
Differing Company Cultures
In his post, Coxon distinguished between the approaches of the two companies he worked for. According to him, at OpenAI employees have not fully internalized the civilizational stakes of the game underway, while at Anthropic the stakes are well understood, but the company feels trapped in a race to be first to achieve breakthrough capabilities.
As possible solutions, Coxon pointed to coordination among industry companies and temporary moratoriums on developing the most advanced model capabilities until credible methods for ensuring their safety exist.
What It Means for the Industry
The posts from Coxon and Hubinger triggered a wide response in tech and financial media, partly because they came at a moment when Anthropic is preparing for investor talks ahead of a potential public offering and pitching the AI market as worth tens of trillions of dollars. A senior employee's public admission that the company has no plan for a core technological risk raises questions about how that narrative squares with the pace of commercialization.
Reactions on social media have been mixed, ranging from praise for both researchers' candor to skepticism about whether such statements will actually change the business decisions of companies investing billions of dollars in developing ever more powerful models.
