Back to Discover

Anthropic agents formalize Fermat's Last Theorem in Lean in 11 days

Official Anthropic card art for Formalizing Fermat's Last Theorem

On Sept. 4, 2026, Anthropic said Claude agents produced the first end-to-end Lean formalization of Fermat's Last Theorem in about 11 days, writing roughly 13 million lines of code. Kevin Buzzard compiled the repo and said it is autoformalization of a known 1990s argument, not new math; Nature covered the milestone on Sept. 7.

Anthropic said on Sept. 4, 2026, that Claude agents produced the first end-to-end, computer-checked Lean formalization of Fermat's Last Theorem in about 11 days. Nature covered the announcement on Sept. 7 and framed it as machine verification of a known human proof, not a new solution of the theorem.

Anthropic reported that the campaign ran on Prove2Me, an open collaborative formalization platform designed by Tianyi Peng and collaborators at Columbia University. The finished artifact, Anthropic said, contains about 13 million lines of Lean, proves roughly 29,500 intermediate theorems used in the final chain, and consumed about six billion output tokens from a general-purpose internal research model Anthropic described as roughly comparable to Claude Fable 5.1.

The company said early multi-agent attempts stalled when agents lost project state, and that Prove2Me's directed theorem graph, separated statement and proof files, and natural-language theorem search made parallel work tractable. Anthropic says the Lean proof uses only Lean's three standard axioms, that a comparator confirmed the theorem statement matches Mathlib's statement of Fermat's Last Theorem, and that human mathematical input was limited to occasional high-level steering from Peng.

Mathematically, Anthropic says the argument follows a simplified Darmon-Diamond-Taylor exposition of the Wiles-Taylor-Wiles proof rather than the newer formalization path Imperial College London's Kevin Buzzard has been building. Anthropic contrasts the project with AI work that claims novel mathematics: here the novelty is automated formal verification of an existing proof.

Buzzard independently compiled the repository and ran Lean's comparator, writing on his Xena Project blog the same day that the code base checks out. He congratulated the autoformalization milestone, stressed that the work adds essentially no new mathematics, noted the Darmon-Diamond-Taylor route rather than his modern blueprint, and said his EPSRC project still aims at Mathlib pull requests and human-readable exploration tools the company release does not replace. Nature separately quoted Buzzard saying that two years earlier such a formalization would have been a fantasy, and Rutgers number theorist Alex Kontorovich saying the 13-million-line computer-checked artifact "just completely blew my mind."

What remains open is how much of the 13-million-line artifact can be refactored into maintainable Mathlib contributions, how reviewers will police Lean soundness edge cases at this scale, and whether the same Prove2Me-style swarm can keep pace with live research literature rather than a settled 1990s exposition. Anthropic's token and line counts are company measurements; Buzzard's compile-and-comparator check is the strongest outside technical read so far.

  • @BinaryVerseAI Video

    Outsider walkthrough of what Claude's Lean FLT formalization did and did not prove (further-watching; not primary evidence).

    Watch on YouTube
  • @IntellectuallyCurious Video

    Intellectually Curious Podcast episode on Claude's autonomous FLT formalization.

    Watch on YouTube
  • @Standarity Video

    Standarity read-through of Anthropic's FLT formalization write-up.

    Watch on YouTube