Artificial intelligence has spent the past few years demonstrating that it can write essays, generate images, produce software and answer complicated questions.

A more consequential test is beginning to emerge:

Can AI contribute to mathematics that humans themselves find difficult?

Recent developments suggest the answer may increasingly be yes.

Anthropic says its Claude AI system has produced the first complete computer-checked formalisation of Fermat’s Last Theorem, working largely autonomously for 11 days to translate the famous mathematical proof into the Lean programming language.

Meanwhile, OpenAI has announced separate research involving an AI-generated proposed solution to the Navier–Stokes problem.

The developments do not mean AI has suddenly replaced mathematicians.

But they point toward a potentially important transition from AI systems that reproduce known knowledge toward systems that can assist with—or eventually contribute to—research.

Why Fermat’s Last Theorem Matters

Fermat’s Last Theorem is one of mathematics’ most famous problems.

In simple terms, it states that there are no positive whole-number solutions to:

aⁿ + bⁿ = cⁿ

when n is greater than 2.

Pierre de Fermat wrote the claim in the 17th century and famously suggested he had a proof that would not fit in the margin where he wrote it.

No such proof was ever found.

The problem resisted mathematicians for centuries until Andrew Wiles developed a successful proof in the 1990s.

The proof relied on sophisticated modern mathematics far beyond what existed during Fermat’s lifetime.

Claude Did Something Different

Claude did not independently discover Wiles’s original mathematical result.

Instead, Anthropic used the model to formalise the proof.

Formalisation converts mathematical reasoning into a language that a computer proof assistant can verify step by step.

In this case, that language was Lean.

This matters because ordinary mathematical proofs are written for humans.

Even expert reviewers can miss gaps.

A formal proof is constructed so that a computer can algorithmically check whether every logical step follows correctly.

Anthropic says Claude worked largely autonomously for 11 days to produce the formalisation.

Why Formalisation Is Difficult

Turning a human mathematical argument into computer-checkable logic is painstaking work.

Humans routinely skip steps that seem obvious.

Computers cannot.

Every assumption and logical transition has to be expressed precisely enough for the proof assistant to verify.

That makes formalisation time-consuming even when the underlying theorem has already been proven.

AI could potentially accelerate this process dramatically.

From Verification to Discovery

Formalising known mathematics is one thing.

Producing genuinely new mathematics is much harder.

That is why the next stage of AI research is so important.

Researchers want to know whether advanced models can identify useful conjectures, discover new relationships and construct original proofs.

OpenAI’s recent work involving the Navier–Stokes problem adds to that discussion.

Navier–Stokes equations describe how fluids move and underpin areas ranging from weather modelling to aircraft design.

Questions surrounding the equations are so important that they form one of the Clay Mathematics Institute’s Millennium Prize Problems.

Any claimed progress in this area requires extremely careful independent verification.

That distinction is essential.

AI producing a plausible mathematical argument is not equivalent to the mathematical community accepting a proof.

AI as a Research Partner

The more immediate opportunity may be less dramatic but enormously useful.

AI could become a research assistant for mathematicians.

Imagine a system capable of:

Searching enormous mathematical literatures.

Formalising proofs.

Checking intermediate steps.

Testing conjectures.

Writing computer code.

Exploring thousands of possible approaches.

Finding connections between different fields.

The human researcher would still determine which problems matter and evaluate the significance of results.

But AI could dramatically expand the number of directions researchers can investigate.

Why Mathematics Is a Powerful AI Test

Mathematics has one major advantage over many other fields:

Answers can often be verified rigorously.

An AI-generated essay may sound convincing while containing subtle errors.

A formal mathematical proof can potentially be checked by software.

That creates an unusually powerful environment for testing AI reasoning.

The system can generate an answer.

Another system can verify whether the logic actually works.

This combination of generation and verification could become important well beyond mathematics.

The Bigger Picture

Scientific discovery is one of the most important potential applications of advanced AI.

If models can help researchers solve difficult mathematical problems, similar systems might eventually accelerate work in physics, chemistry, biology and engineering.

Anthropic is already exploring applications in areas including protein design, chemistry and bioinformatics.

Other AI developers are pursuing similar scientific applications.

The long-term promise is not simply an AI that knows what humanity already knows.

It is an AI that helps humanity discover what it doesn’t know yet.

What Happens Next

Claims of AI-generated scientific breakthroughs should be treated carefully.

Extraordinary results require independent verification.

Researchers also need to distinguish between systems retrieving or recombining existing ideas and systems producing genuinely novel contributions.

But Claude’s formalisation work demonstrates something concrete.

AI can already perform parts of sophisticated mathematical workflows that previously required enormous amounts of specialist human effort.

The next question is much bigger.

Can artificial intelligence move from checking the mathematics we know to helping discover the mathematics we don’t?