News

Claude formalizes Fermat Last Theorem in Lean

Anthropic's research post says Claude produced the first Lean-checked proof of Fermat's Last Theorem in 11 days. It wrote 13 million lines and about 29,500 intermediate theorems. The run used about six billion output tokens on an internal model comparable to Claude Fable 5.1. This formalizes Wiles, it is not a new proof.

Claude formalizes Fermat Last Theorem in Lean

Anthropic says Claude produced the first complete computer-checked proof of Fermat's Last Theorem, working largely autonomously in Lean over 11 days. The research post (4 September 2026) is a lab blog with Lean artifacts on GitHub, not a peer-reviewed journal paper.

This is autoformalization of a known proof. Claude followed a simplified Wiles path from Darmon, Diamond and Taylor. It did not discover a new proof of FLT, and it did not recover Fermat's margin argument.

Along the way Claude wrote 13 million lines of Lean and proved about 29,500 intermediate theorems in the final artifact (about 30,300 along the way). Human mathematical input was limited to occasional high-level notes from Tianyi Peng. Early multi-agent attempts stalled. The run succeeded after the team switched to Prove2Me, Peng's collaborative formalization scaffold.

The token bill is the part most headlines skip. Anthropic says the campaign consumed about six billion output tokens from an internal research model roughly comparable to Claude Fable 5.1. Failed early collaboration still left about 7% of the non-boilerplate lines in the finished proof. Lean checked the artifact against Lean's three standard axioms, and a comparator confirmed the statement matches Mathlib's own FLT statement.

Kevin Buzzard reviewed the result and said it proves FLT with no assumptions other than the axioms of mathematics. He also said autoformalization artifacts are now robust enough to build on. Anthropic is explicit that a Lean check does not replace a human-readable exposition.

The same lab is already selling enterprise controls around long-running agents, including frontier safeguards and retention. Unattended agent swarms have a different failure mode, the one in OpenAI agents writing to a German wiki. Formalization at this scale also sits next to the model-hub deal tape, including Nvidia's agreed Hugging Face purchase and Google's Gemini agentic video work.

If you run a research lab, a formal-methods product, or diligence on AI math claims this week, treat Lean plus a Prove2Me-style scaffold as the product surface, not a raw chat window. Require an axiom-only Lean check and a Mathlib-statement comparator before you cite an AI formalization in a paper or a memo.

Subscribe to Techpresso

Free daily newsletter, read in 5 minutes.

Subscribe free