What long-horizon coding agents did to mathematicians, and what RSEs can learn from the incident and from how they reacted
Research Software & Analytics Group, University of Exeter
September 30th, 2026
Mathematics is the canary in a coal mine for all parts of intellectual life. Its claims can be carefully checked and so we are witnessing the disruption in real time. […] Science, law, public policy, and art will all have to find ways to assess what is valuable amidst the abundant output.
— Bryna Kra, Deep theorems were scarce, September 2026
Research software included. See my link post.
The limit
Verifiable signals, and Lean: a proof checked by a machine.
The incident
From the IMO to the Jacobian conjecture and Navier–Stokes.
The reaction
What mathematicians say they are after. And a parallel in cybersecurity.
Dialling back
What changes for research software, on two axes.
Then discussion, and a list of questions I don’t have an answer to yet.
mathlib library produce the terms that it checks.Software has formal verification too, e.g. seL4, CompCert, and Amazon’s use of TLA+. We will come back to why it stays rare.
I think this may be my most important theorem to date. […] nobody else has dared to look at the details of this […] Better be sure it’s correct…
Nothing to do with advancing AI. Both stories are in Hartnett (2026). Scholze’s challenge is on the Xena blog (Scholze 2020).
July 2024
AlphaProof + AlphaGeometry 2: IMO silver. Problems manually translated into Lean; up to three days.
July 2025
Gemini Deep Think and an OpenAI model: IMO gold, in natural language, within the 4.5 hours.
July 2026
A counterexample to the Jacobian conjecture, found with Claude Fable 5.
September 2026
OpenAI: Navier–Stokes blowup. ~10,000 agents, 88 h search + 17 h Lean, 165 pages.
In 2024, Lean was what let the model do it at all. By 2026, the proof is found in natural language, and Lean is the certificate at the end.
\begin{aligned} F(z_1,z_2,z_3) = \big(&(1+z_1 z_2)^3 z_3 + z_2^2 (1+z_1z_2) (4+3z_1z_2), \\ &z_2 + 3 z_1 (1+z_1z_2)^2 z_3 + 3 z_1 z_2^2 (4+3z_1z_2), \\ &2 z_1 - 3 z_1^2 z_2 - z_1^3 z_3\big) \end{aligned}
Easy to check, and it looks like a miracle. Tao still wrote a whole post digesting it. The question was asked by Akhil Mathew.
OpenAI says it doesn’t intend to claim the prize. Clay’s rules need publication, two years, and general acceptance by the community.
\begin{aligned} \partial_tu+(u\cdot\nabla)u&=-\nabla p+\nu\Delta u+f,\\ \nabla\cdot u&=0. \end{aligned}

For the mathematics, see Unpacking the Navier–Stokes blowup. For why I call it an incident, see the series guide.
| no force, f=0 | with a force, f\neq0 | |
|---|---|---|
| Euler, \nu=0 | blowup claimed by OpenAI; a separate candidate from Anandkumar’s group | blowup claimed by Alpöge and Buckmaster |
| Navier–Stokes, \nu>0 | still open (Clay’s A/B) | blowup claimed by OpenAI (Clay’s C/D) |
Claims as of 11 September 2026. See the four corners.
Like AlphaGo
Not like AlphaGo
It doesn’t need to be an AlphaGo moment to disrupt how mathematics trains people, assigns credit, and supports careers.
divergence the classical divergence? Does smooth include t=0? A mistake here gives a verified proof of a different theorem.sorry, no custom axiom, no native_decide. #print axioms should list only propext, Classical.choice and Quot.sound.Perhaps a day or two of expert attention, instead of a year of refereeing. Details in the Lean section.
Mathematicians aren’t mourning their craft. See After Math, translated for software engineers.
A Severe Misalignment of AI in Mathematics, 11 September 2026:
And from the essays that followed: journals should “distinguish and credit the roles of discovery, proof, formalization and explanation” (Kra 2026).
We will come back to this list in the discussion: what is our equivalent of each, and do we uphold it?
Timothy Gowers didn’t sign, because of the “tool and proxy” sentence:
Remember the problem-solving end of that spectrum. See A flood from within the community.
| Cybersecurity | Mathematics | |
|---|---|---|
| Persistence, parallelism | Mythos Preview chains four vulnerabilities into one browser exploit | 10,000 agents over 88 hours |
| Hard to digest | curl’s security reports arriving 4–5× faster than in 2024 | a Lean-verified proof still has to be understood |
| A rumour is enough | probes about ten minutes after a public fix | OpenAI started its search on a rumour, which was wrong |
| The rush lands on people | Daniel Stenberg and the curl team | Buckmaster, Anandkumar: unfinished work rushed out |
Security answered with coordinated disclosure. My proposal for mathematics: progressive disclosure.
Correctness
Understanding
Not people versus machines as such: the four-colour theorem and Hales’ Kepler proof were monoliths made by people, and Mochizuki’s layers are ones nobody else shares.
Researcher
What is in their head, and what they say.
RSE
Listening and asking: user stories into requirements. And back again: training.
LLM
A lower-level translator: the spec into a program.
Program
Checked against the oracle, where there is one.
Like checking an English theorem against its Lean statement, the human work is at the boundary. Unlike Lean, there is no kernel after it.
I haven’t solved this. See Born dead.
| Mathematicians uphold | Our equivalent? Do we uphold it? |
|---|---|
| Understanding is the goal (Fields medallists) | communication? |
| Time for a proper writeup (Fields medallists) | documentation that explains why? |
| Credit for formalization and explanation (Kra) | credit for tests, documentation, maintenance? |
| A solution that opens lines of work for others (Cohn) | open source projects? knowledge exchange? |
| Students and ideas (Fields medallists, Gowers) | handover, training? |
Discussion
Slides and the posts behind them: blog.kolen.dev