Back

Anthropic researcher Jacob Coxon resigns, warning labs are gambling with AI risk

Anthropic wordmark still used for resignation-week coverage

Former Anthropic pretraining researcher Jacob Coxon said on Sept. 9, 2026, that he resigned because Anthropic and OpenAI are racing toward self-improving superintelligence. Anthropic Alignment Science Lead Evan Hubinger publicly agreed that staff earnestly believe AI could kill all humans, and an Anthropic spokesperson told Wired the industry should adopt a lawful way to pace powerful model releases.

Jacob Coxon, a pretraining researcher who said he had spent the last three years at OpenAI and Anthropic, announced on Sept. 9, 2026, that he resigned from Anthropic. In a public thread he said neither company is acting responsibly and that both are "racing straight to self-improving superintelligence and gambling with our lives."

Coxon wrote that people building AI earnestly believe it could kill everyone by the end of the decade, that the danger is not a marketing stunt, and that warning shots such as the Hugging Face agent incident make pacing agreements among U.S. labs more viable. He urged lab researchers not to kick off a superintelligent reinforcement-learning run without a rigorous understanding of the system.

Independent: In an Axios interview the same week, Coxon said he left about two months before his Anthropic equity would have vested, after roughly four months at the company, and that he still holds equity in OpenAI. He told Axios he has not seen Anthropic compromise safety to outlast rivals so far, but feared race pressure would force cut corners later.

Anthropic Alignment Science Lead Evan Hubinger replied publicly that Coxon is correct and that staff "earnestly believe AI could kill all humans." Hubinger gave a personal estimate of greater than 10% within the next decade, said Anthropic is trying its best, and said the company does not yet have a plan to solve alignment for superintelligence and is not clearly on track. Treat the percentage as his personal claim, not a FrameSignal fact.

An Anthropic spokesperson told Wired the company has been transparent that AI brings benefits and unprecedented risks, cited safety work such as mechanistic interpretability, and said the world would benefit from a lawful, verifiable way for the industry to pace how it releases powerful models. OpenAI did not return Wired's request for comment at the time of that article.

Anthropic's Institute page "When AI builds itself" says the company is already delegating a growing share of AI development to AI systems, that full recursive self-improvement is not inevitable, and that the trend could arrive sooner than most institutions are prepared for. Anthropic's redacted August 2026 Risk Report is the company's formal Responsible Scaling Policy disclosure covering autonomy and misalignment threat models for covered models.

A July 2026 open letter at pacingthefrontier.com, signed by 1,386 frontier-lab employees, asked the U.S. government to help build tools to deliberately pace automated AI development. Three days after Coxon's posts, Anthropic CEO Dario Amodei published "We Must Pace the Frontier." That Sept. 12 CEO plan is covered in a related story. OpenAI's Sept. 9 "The AI policy window is open" post is adjacent policy context.

  • @t3dotgg Video

    A YouTube walkthrough of the same-week pacing debate, including the resignation posts and lab replies.

    Watch on YouTube