The flood, released responsibly?

Math
LLM
Research ethics
OpenAI released 722 AI-generated manuscripts after consulting the mathematicians’ advisory group. An improvement on the Navier–Stokes incident, short of what the group asked, and the flood Gowers predicted.
Author

Kolen Cheung

Published

October 8th, 2026

On 6 October, OpenAI announced a release of AI-generated mathematics, after consulting the mathematicians’ advisory group:

We’re releasing a broad range of new mathematical results produced by an internal frontier model.

As we look to improve how we share results with the math community, we’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study to develop best practices, and we have drawn on their advice and public recommendations to inform how we release these results.

[…] The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.

The results are in a GitHub repository, whose README says how many there are, and why:

As part of model development, we evaluate our models on open research problems. We expanded these evaluations after performance on our existing mathematical evaluations saturated. […]

The current catalogue contains 722 manuscripts organized into 372 families. […] Over the course of the evaluation, the model was posed approximately 4,000 problems.

This is the release the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) was advising on when it said it was helping OpenAI “coordinate the release of a large number of significant results”, as I quoted at the end of What if the button came with the understanding? A week before the release, the group published its guidelines for AI labs, and they open with:

At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.

If AI labs produce significant mathematical results, they should responsibly release the results, as outlined in this document, as soon as possible.

AI labs that release substantial mathematical output without immediate accompanying human understanding must take responsibility for ensuring that human understanding will follow.

Here is what the guidelines ask of a lab releasing results that nobody understands yet, against what OpenAI has done, as of 8 October 2026:

AGMAI asks OpenAI’s release
Stop testing advanced problems on proprietary models. No: “This is why it is important to continue to evaluate our internal frontier models on mathematics and other sciences”.
Search the literature, and cite the papers where the ideas were first introduced. Partly. The manuscripts have reference lists, and OpenAI commits to “improving the quality of the papers via the citations, mathematical exposition, and presentation” for future releases.
Write each proof up as a traditional mathematical paper. Partly, under the same commitment.
Deposit the results in a repository no AI lab controls, with persistent citable identifiers, recorded revisions and comments. Mostly not. It is a GitHub repository under OpenAI’s own organization, with issues turned off, and a BibTeX entry but no DOI. Revisions will be kept as new versions, and OpenAI is “continuing to explore other community-hosted alternatives”.
Don’t treat the release as a marketing vehicle. No. It was announced on OpenAI’s own blog.
For each result: the model name, the prompts, a summarized chain of thought, the time taken and the cost. Partly. “An unreleased internal OpenAI model”, no prompts, reasoning summaries for 10 of the 372 families, and the compute as an average over all results.
Formalize, with a challenge file for Comparator and a formalization.yaml. Mostly. “Many, but not all, of the manuscripts have been formalized” in Lean, in that form.
Say how AI came to be used on each problem, how many comparable problems it failed on, and how the problems were chosen. Partly. About 4,000 problems posed, filtered by “requiring an appropriate level of significance”, but not how they were chosen.
Fund the work of understanding, with existing nonprofits, not the lab, deciding where the money goes. Promised: “We will be funding a series of workshops, conferences, and special programs”. Who decides is not said.
Broad, equitable access to the models. Promised for this one: OpenAI is “working to responsibly release” it.

AGMAI’s statement on the day of the release leaves the grading to us:

AGMAI’s advisory role should not be interpreted as a judgment of the impact of these results or an endorsement of the process by which OpenAI obtained them. […] While we consider these discussions constructive, it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully, and whether there are others we should suggest.

My tentative read is that it is an improvement on the Navier–Stokes incident, but it isn’t fully what AGMAI asked. OpenAI calls it a disclosure, and promises to “update our standards for future disclosures of major scientific advancements”. It is a coordinated one this time, with AGMAI consulted beforehand, but not a progressive one: 722 manuscripts landed on one day. Ideally, if it were up to AGMAI, these results would not exist yet, because they come from exactly the kind of proprietary model it asked the labs to stop testing.

This is the flood, as Gowers prophesied:

[…] there will be a flood of new results, whether we like it or not, and it will no longer be the AI companies producing them, though perhaps the pattern will continue that the AI companies will have access to more powerful models and so will obtain more than their fair share of headline results. (Gowers 2026)

If they hold up, it is also, for the first time (except maybe for one occasion earlier this year (Brenner et al. 2026)), a major contribution to mathematical physics by AI math.1 Three of the ten reasoning summaries are mathematical physics: the Mézard–Parisi formula for diluted spin glasses, spontaneous magnetization in the quantum Heisenberg ferromagnet, and the three-dimensional relativistic Vlasov–Maxwell system.

One way to see that this is a flood is the cost. On average, a result is equivalent to three hours of ChatGPT Pro’s work, so it is no longer Navier–Stokes level costs, which were about $6.5 million at public API prices for that one problem. It is affordable enough that every researcher, if they want to, can be equipped, once the model is released.2

It is also the scenario where this tsunami is more difficult to stop. If it comes in waves, once per new model, the disruption may be containable by coordination, which is what progressive disclosure was about. But if every ChatGPT Pro user can achieve this, then whether OpenAI wants to stop or not is irrelevant. Once it is put in the hands of individuals, some people will use it and create the same flood, as Bloom has already seen on a smaller scale with AI proofs on the Erdős problems site. So perhaps it makes the responsibility of OpenAI, and by extension the AI labs, lesser, because it is no longer an ethical issue of the AI labs, but humanity’s. Can you convince each individual to listen to the mathematicians, behave themselves, and not push the button? I don’t think so. Not that I don’t think this is a problem, but I just don’t believe everyone will behave: the probability of one or more bad apples goes to 1.

OpenAI also makes it explicit that math is a benchmark in this case. Their existing benchmarks are saturated, so open math problems are the one left for them to push the envelope of AI capability. Navier–Stokes was not an AlphaGo moment, and a corollary is that they’ll be back. After AlphaGo, DeepMind released 50 games and moved on to AlphaFold. AI labs will not conquer mathematics and move on. They will continue to use it as a benchmark, which to a certain extent is worse, because (some) mathematicians want to be left alone.

A corollary of math being a benchmark is that the collaboration between AI labs and the math community will continue to be the bare minimum. The majority of the math community doesn’t seem to want to be more closely associated with AI labs. And the AI labs, because the math community lacks interest in a tighter collaboration, and because they keep needing new benchmarks to advance AI capabilities, will continue to do things their own way, with minimal compliance to avoid further conflict. Benchmarking is the priority, and how to benchmark with math while minimizing the harm, perhaps to their image, is a consideration. (I use “AI labs” liberally here. Perhaps I should say OpenAI only, because so far the flood seems to be only theirs.)

Speaking of which, the natural question to ask is, where is everyone else? Google showed one result in March, and Anthropic one from an unreleased model in August (Anthropic 2026), but those aren’t a flood. From the benchmark and marketing perspective, a flood is a clear demonstration of capability. So the question becomes: are the other labs actively not engaging in a flood, or are they less capable? Either OpenAI is far more advanced than everyone else, or OpenAI is far ahead in its willingness to perform demonstrations like this despite the obvious opposition. From OpenAI’s perspective, surely they would want others to think they are leading in capability, not in immorality. Is Anthropic’s internal model less capable? Or does Dario Amodei hold that high a moral ground, not to take part in this kind of thing? Time will tell, when the current generation of internal models falls into individuals’ hands, and they can do their own “three hours of ChatGPT Pro” research. AGMAI, for its part, has “already been in contact with several of them”.

The bottom line: the situation seems to be better than on the day of the Navier–Stokes incident, but it could have been a much tighter collaboration between the two cultures.

References

Anthropic. 2026. “Claude’s Progress on the Riemann Hypothesis.” August 10. https://www.anthropic.com/research/riemann-zeta.
Brenner, Michael P., Vincent Cohen-Addad, and David Woodruff. 2026. “Solving an Open Problem in Theoretical Physics Using AI-Assisted Discovery.” 2603.04735. Preprint, arXiv, March 5. https://doi.org/10.48550/arXiv.2603.04735.
Gowers, Timothy. 2026. “Why I Didn’t Sign the Fields Medallists’ Letter.” Gowers’s Weblog, September 17. https://gowers.wordpress.com/2026/09/17/why-i-didnt-sign-the-fields-medallists-letter/.

Footnotes

  1. By AI math I mean LLMs deriving mathematical results, however the LLM is augmented, excluding other ML approaches unrelated to LLMs.↩︎

  2. That is three hours per result, not per problem attempted: the model was posed about 4,000 problems for 372 families. And it is an unreleased model, not the one ChatGPT Pro users have today.↩︎