TL;DR — Key Takeaways

  • Anthropic says Claude agents produced the first complete computer-checked formalization of Fermat’s Last Theorem in 11 days.
  • The project generated 13 million lines of Lean code and used about 29,500 intermediate theorems in the final proof.
  • Anthropic sees AI-assisted formalization as a way to help mathematicians more quickly verify both human and AI-generated work as models take on more research.

Anthropic says a team of Claude agents has produced the first complete, computer-checked formalization of Fermat’s Last Theorem (FLT), the centuries-old mathematical problem first proved by Andrew Wiles and Richard Taylor in the 1990s.

Wiles spent years developing the proof itself, while Claude produced 13 million lines of Lean code for the formalization in just 11 days. Dozens of Claude agents proved more than 30,000 intermediate theorems, about 29,500 of which were used in the final proof.

Formalizing a mathematical proof means translating the argument into a precise form that software can verify step by step, which has historically been painstaking work. Since 2024, mathematician Kevin Buzzard at Imperial College London has led a community effort to formalize FLT in Lean, which was expected to take years. Anthropic said Claude followed a simplified version of Wiles’s proof and built on existing formalized mathematics, including work from Buzzard’s FLT project.

Anthropic shared Claude’s completed formalization with Buzzard for review. He called the result an “extraordinary autoformalization achievement,” saying it showed AI could formalize work spanning algebra, harmonic analysis, geometry and number theory in a form that later mathematical work could build on.

Anthropic’s FLT project initially ran into a problem that may sound familiar to developers building long-running agents. Individual agents made progress but lost track of the larger project and stopped coordinating effectively. Researchers eventually used Prove2Me, an open platform developed by Tianyi Peng, an Anthropic researcher who initiated the FLT experiment, along with collaborators at Columbia University. Prove2Me allowed for tracking how different parts of the proof depended on one another, coordinating work among agents and making completed results easier to find and reuse.

Researchers paired Prove2Me with a Claude Code-based multi-agent workflow, allowing dozens of agents to work on different parts of the proof in parallel. Anthropic said the project consumed about 6 billion output tokens from an internal general-purpose research model with capabilities roughly comparable to Claude Fable 5.1.

Anthropic published the formalization on GitHub and said the proof uses only Lean’s three standard axioms and contains no omitted proofs. That means the proof does not rely on placeholders for unfinished steps or introduce new axioms just to make the argument work. The project also compared the result with Mathlib’s existing statement of Fermat’s Last Theorem and ran it through a second independent Lean kernel implementation.

Anthropic is predicting that AI-assisted formalization could become an important part of mathematics research as models produce more proofs and conjectures. Turning those results into machine-checkable form could help researchers catch errors and reduce the amount of manual review needed. “As formalization becomes a more commonplace tool, we are hopeful that it will help maintain trust in the common body of mathematical knowledge,” Anthropic wrote in a blog post detailing the research.

The company is also putting resources behind that idea. It has expanded free and discounted access, research credits and grants for outside mathematicians working on formalization and related scientific projects, including projects to formalize other major theorems and improve existing software tools. Last month, Anthropic applied a similar formalization process to work by an unreleased Claude model on a problem related to the Riemann hypothesis, another famous conjecture and one of the seven Millennium Prize Problems. If AI ever cracks one of those, formalization will seem like the easy part.

TECHSTRONG AI PODCAST

SHARE THIS STORY