Why math wants the whole chain open

Math
LLM
Research ethics
Open source
Cultural differences
Science tolerates proprietary software as long as the scientific part is open. Math is on the far end of that spectrum, where the end-to-end chain must be open, including how you came up with the answer.
Author

Kolen Cheung

Published

October 10th, 2026

In The flood, released responsibly? I quoted the opening of the guidelines that the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) published for AI labs:

At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. […] we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.

OpenAI’s release came from “an unreleased internal OpenAI model”, and it intends to keep evaluating its internal models on mathematics. But science uses proprietary tools all the time. So is the objection that the models are proprietary, rather than that frontier AI is used on advanced math at all? And why should math be held to a stricter standard than the rest of science?

I think one way to see it is that math, on the spectrum of open vs. closed “dependencies”, is probably on the farthest end: absolute openness. E.g. that’s why you can’t copyright math, as far as I understand.

Even in my field, while our science is open, the software used (a dependency of our science) might be proprietary. E.g. industrial, proprietary, commercial software simulating electromagnetic effects is used, and is core, to design the amazing sinuous antennas of the multichroic detectors in CMB experiments. The ones in Suzuki et al. (2012) were designed with HFSS, a commercial 3D electromagnetic simulator. So part of the design process is not entirely open: the theory behind it is, the software used in the design is not, and the effect of the design is measured in the lab, so it is open again.

This has always been a bit unsettling to me, perhaps to my inner mathematician. But it seems to be normal in science: we use SOTA technology to advance science, and as long as the “scientific” part is open, we tolerate it.

In math, the end-to-end chain “must” be open, and that should include how you came up with the answer as well. So an open model is best, but even a proprietary model should at least be publicly available. Using an internal model means only someone from OpenAI can do that, which is the farthest from that ideal, kind of the opposite of the democratization of math.

This is not entirely new. Durán et al. (2014) were using Mathematica to look for counterexamples to their conjectures generalizing a result of Karlin and Szegő on orthogonal polynomials, which came down to computing determinants of matrices of big integers. Mathematica found some counterexamples:

Fortunately, another of us was using Maple, and when checking those supposed counterexamples he found that they were not counterexamples at all.

Mathematica was computing the determinants wrongly, and gave a different answer when the same determinant was evaluated twice. Maple and Sage agreed on the right one. And their conclusion is the same unsettling feeling:

The commercial computer algebra systems are black boxes and their algorithms are opaque to the users (of course, also the source code), and certainly this does not contribute to avoid errors. […] Moreover, known bugs of computer algebra systems should be available to the users; this is usual in free software, but an anathema for commercial packages.

Using more than one system is a way to reduce the trust in a single proprietary software, and interestingly, comparisons like this sometimes uncover hidden bugs. But that only reduces it. Again, the scale and reach of LLMs make that issue more prominent and intolerable, and with an internal model there isn’t even a second system to check against. AGI, where the G is general, cuts both ways: benefits and harms are both generally available (and you’re welcome).

References

Durán, Antonio J., Mario Pérez, and Juan L. Varona. 2014. “The Misfortunes of a Trio of Mathematicians Using Computer Algebra Systems. Can We Trust in Them?” Notices of the American Mathematical Society 61 (10): 1249–52. https://doi.org/10.1090/noti1173.
Suzuki, Aritoki, Kam Arnold, Jennifer Edwards, et al. 2012. “Multichroic Dual-Polarization Bolometric Detectors for Studies of the Cosmic Microwave Background.” Millimeter, Submillimeter, and Far-Infrared Detectors and Instrumentation for Astronomy VI, Proceedings of SPIE, vol. 8452: 84523H. https://doi.org/10.1117/12.924869.