The economy of AI math, for mathematicians

Math
LLM
Research ethics
Cultural differences
Why would OpenAI keep testing open problems on its internal models, when the mathematicians’ advisory group asked it to stop? My guess: because of how an LLM is paid for.
Author

Kolen Cheung

Published

October 9th, 2026

OpenAI has released 722 AI-generated manuscripts after consulting the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), whose guidelines “ask them to stop testing advanced mathematical problems on proprietary models”. OpenAI’s answer is that “it is important to continue to evaluate our internal frontier models on mathematics and other sciences”. Here is what I think causes the fundamental tension between OpenAI and AGMAI: the economy of building an LLM.

This is speculation. I have no internal knowledge of OpenAI, and what follows is built from publicly available sources only. As far as I can tell it is consistent with them, but that doesn’t mean it is the only possibility.

Training a new LLM is expensive. Nobody outside the labs knows exactly how expensive, so here is a range of best guesses. At an MIT event in 2023, Sam Altman “was asked if training GPT-4 cost $100 million; he replied, ‘It’s more than that.’” (Knight 2023) Epoch AI estimates the final training run of xAI’s Grok 4 at about $490 million, and that the cost of the largest runs has grown 2.4 times a year since 2016, which would put them above a billion dollars by 2027 (Cottier et al. 2025). Dario Amodei expected it sooner: in 2024 he said models costing closer to $1 billion were already “in training today”, and that training costs would reach $10 billion and $100 billion in “2025, 2026, maybe 2027”.1

Inference is where they make money. As Sam Altman told reporters in August 2025:

We’re profitable on inference. If we didn’t pay for training, we’d be a very profitable company.

I can’t check whether that is true, but this is how OpenAI sees its own economy, and that is what drives its incentives.

Training is also where the capability comes from, and today’s SOTA LLM’s biggest weapon is reinforcement learning with verifiable rewards (RLVR): the model attempts a task whose answer can be checked automatically, the tests pass or the proof compiles, and is rewarded when it does. Andrej Karpathy, looking back at 2025, called it the new major stage of training, and said “most of the capability progress of 2025 was defined by the LLM labs chewing through the overhang of this new stage” (Karpathy 2025). And it is how a model can move into another class so quickly: code is among the most verifiable tasks, as the tests pass or fail, and from the coding gains came the cybersecurity breakthroughs of the Mythos moment, which according to Anthropic “emerged as a downstream consequence of general improvements in code, reasoning, and autonomy” (Carlini et al. 2026). The recent advances of AI math are the same story, with a checkable answer or Lean as the verifier.

RLVR and test-time compute look similar: an agent working in a loop, getting a verifiable signal so that it can “converge” to the right solution. The difference is that RLVR improves the model, and the latter doesn’t. Whatever the model figures out at test time lives in its context, and is gone when the session ends. Keeping what it learnt in the model is called continual learning, which is still not solved.2 So the same math problem is worth something different to OpenAI depending on when it is run. At train time it can improve the model, and at test time it only produces an answer.

This is where the usual assertion surrounding their economy comes in, namely that they spend so much compute on math, costing so much money. Álvaro Lozano-Robledo, in a guest post on Terence Tao’s blog, asks it as a question of what society should pay:

Should we spend millions of dollars and an undisclosed amount of natural resources in order to find a solution for Navier-Stokes? […] we may reach a point where an LLM could solve an important problem for an exorbitant cost (in terms of funding and resources) but it may just not be an acceptable cost for the taxpayer or society to bear. (Lozano-Robledo 2026)

Even if math is just a benchmark, I think the assertion is (probably) wrong, as far as OpenAI is concerned. Priced at public API rates, as I did for Navier–Stokes, all the problems OpenAI attempted for that announcement, Navier–Stokes included, came to about $15 million, and that is what the compute would cost us, which is more than it costs OpenAI. Against the hundreds of millions to billions for a training run, that is spare change.3 And evaluating is part of model development: a model is evaluated while it is trained, before it is released. Once it is released, the evaluation has done its job, and the incentive to spend even this spare change disappears.

And if it is also RLVR, using math as a benchmark for an internal model is itself valuable: it gives value rather than (just) consuming it. Yes, this is a big assumption, because OpenAI calls it evaluation, in the README of the repository:

As part of model development, we evaluate our models on open research problems. We expanded these evaluations after performance on our existing mathematical evaluations saturated.

Normally you don’t train on your evaluation set, or it stops measuring anything. I think open problems are different, because they don’t run out: what the model solves can go into training the next model, as in STaR, where a model is trained on its own answers that turned out correct (Zelikman et al. 2022), and what it doesn’t solve stays an evaluation. There is a verifier: “Many, but not all, of the manuscripts have been formalized” in Lean. And the success rate is about right for RL: 372 families of results from about 4,000 problems, roughly one in ten. RL learns little from a problem the model always solves or never solves, and “saturated” is, in RL terms, a curriculum that ran out of signal. I haven’t found anything public that says this is not what happens.

In that case these runs are not (merely) consuming but providing signal, and the marginal cost of the math results is close to zero.4

What about the natural resources Lozano-Robledo also asks about? I don’t think this marginal spend is the place to count them. If we are really concerned about the environment, the cost should be compared to the value the whole model delivers in its lifetime, because the benchmark is part of the cost of the model. The illusion is that you could train an as strong model without doing this benchmark. If it is RLVR, you literally couldn’t, as the runs are training. If it is only a benchmark, you could not steer towards as strong a model without it: a saturated evaluation can’t tell you which checkpoint or recipe is better. Either way, given how the labs choose to build models, the model and the benchmark are inseparable, and the question moves from “should we spend this to solve Navier–Stokes?” to “is the whole model worth what it costs?” That is a question about frontier AI in general, and a much bigger one than this post.

And therefore, whether it is a benchmark or RLVR, the only point where this economy works is at train time, i.e. while the model is still internal, which is in tension with what AGMAI is advising. And there is always a model at train time: the frontier is always an internal model in training, one generation ahead of the public one. So I think they’ll be back.

To reiterate, they have no incentive to use their publicly available model to produce these kinds of results: it would cost them money, in compute they could have sold, for a model that would learn nothing from it. That is the (test) time to sell subscriptions, and let others use their models to find new results. Their own announcement measures the work in a unit of their product, “roughly three hours of ChatGPT Pro thinking”.5

So fundamentally, what the math community, represented by AGMAI, is asking of the AI labs (OpenAI specifically) is the opposite of their best interests. To stop testing advanced problems on proprietary models is to give up the problems they need at the frontier of model development.

My guess is that OpenAI’s train of thought is more like this: we created these new math results as a by-product of using math as a benchmark, and now that we have them, should we release them? According to AGMAI, the answer is an emphatic yes:

If AI labs produce significant mathematical results, they should responsibly release the results, as outlined in this document, as soon as possible.

Even if OpenAI wanted to abide by the advice not to test, it was too late.6 The results had already been found by the internal model before OpenAI asked.

To make the point even more plainly: they are not asking whether they can use math as a benchmark. They believe it is their god-given right (or god-making right, depending on who you’re asking) to use math any way they want. What they are asking is, given that’s what they will do, how else should they be doing it? They are looking for a win-win given math as a benchmark, and they are not willing to remove that “prior”.

If this is right, it may also answer my question last time, where is everyone else? Any lab doing RLVR has the same incentive to test on open problems internally, so the difference may be that OpenAI released its results and the others didn’t. Is that primarily because OpenAI is abiding by AGMAI’s request to release as soon as possible? Or is it purely a marketing stunt, since more results make them look more capable?

Whether or not the story I’m telling is exactly true, I hope mathematicians can put on the shoes of an AI researcher and see why their research culture has such a different incentive to use math. And unfortunately (or maybe fortunately), math belongs to all humanity, so everyone can use it, much as open source comes with the right to use it and no warranty, and see its beauty from their own point of view, even when those points of view sometimes look so alien to one another.

References

Carlini, Nicholas, Newton Cheng, Keane Lucas, et al. 2026. “Assessing Claude Mythos Preview’s Cybersecurity Capabilities.” Anthropic, April 7. https://www.anthropic.com/news/mythos-preview.
Cottier, Ben, Robi Rahman, Loredana Fattorini, Nestor Maslej, Tamay Besiroglu, and David Owen. 2025. “The Rising Costs of Training Frontier AI Models.” arXiv:2405.21015. Preprint, arXiv, February 7. https://doi.org/10.48550/arXiv.2405.21015.
Hubert, Thomas, Rishi Mehta, Laurent Sartran, et al. 2025. “Olympiad-Level Formal Mathematical Reasoning with Reinforcement Learning.” Nature 651 (8106): 607–13. https://doi.org/10.1038/s41586-025-09833-y.
Karpathy, Andrej. 2025. “2025 LLM Year in Review.” Karpathy, December 19. https://karpathy.bearblog.dev/year-in-review-2025/.
Knight, Will. 2023. “OpenAI’s CEO Says the Age of Giant AI Models Is Already Over.” Wired, April 17. https://www.wired.com/story/openai-ceo-sam-altman-the-age-of-giant-ai-models-is-already-over/.
Lozano-Robledo, Álvaro. 2026. “What Should We Tell Our Students?” What’s New, October 8. https://terrytao.wordpress.com/2026/10/08/what-should-we-tell-our-students/.
Zelikman, Eric, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. 2022. “STaR: Bootstrapping Reasoning with Reasoning.” Advances in Neural Information Processing Systems 35. https://doi.org/10.48550/arXiv.2203.14465.

Footnotes

  1. These are not quite the same quantity. Epoch counts the hardware and energy of the final run only, which is why its own estimate for GPT-4 is about $40 million (Cottier et al. 2025), well below Altman’s. Staff, experiments and failed runs come on top, and Epoch found staff alone to be 29% to 49% of the total for GPT-4 and Gemini Ultra.↩︎

  2. There is an exception. For the hardest problems, AlphaProof ran RL at test time on generated variants of the problem it was solving (Hubert et al. 2025), so its weights did change at test time. But that adaptation was for one problem and was not kept, so it isn’t continual learning either.↩︎

  3. An order of magnitude only: the training figures are for other models, some from other labs. But even several times more would still be small next to training.↩︎

  4. Perhaps even negative, if the signal is worth more than the compute it took, but nobody outside OpenAI can know that.↩︎

  5. Public use isn’t worth nothing to them either. By default, conversations with ChatGPT may be used to train their models unless the user opts out. But the internal loop has the choice of problems and a verifier, and a subscriber’s session has neither.↩︎

  6. AGMAI itself was set up after OpenAI approached some of its members, and its first task was advising OpenAI on “how to coordinate the release of a large number of significant results in mathematics that they report have been produced by their internal model”, as I quoted in What if the button came with the understanding?↩︎