Can a theory be laid out, and can it be rewarded?
In Long-horizon agents write dead programs I raised two questions in passing: can a theory be laid out at all, and can it be rewarded? Logically, I think the second depends on the first. If we can’t lay out the theory, how could we reward it? And I only have an answer to the first in a restricted domain, science. So this post is more like an open question, exploring an answer rather than giving one.
Can a theory be laid out?
Naur would say no. For him the program is the theory in its programmers’ heads, and the knowledge they have “necessarily, and in an essential manner, transcends that which is recorded in the documented products”. It depends on grasping which situations in the world are similar, which is why that knowledge “could not, in principle, be expressed in terms of rules.” (Naur 1985)
Emily Riehl’s theory building points the other way. Its aim is to “make it easier for more people to hold increasingly complicated mathematical thoughts in their heads”, by finding a clarifying abstraction at the “right” level of generality (Riehl 2026) (see my link post). And that abstraction does get written down, in the definitions and the papers.
Are they in tension, or two halves of one thing? One way I see it is that abstractions are the expressible skeleton of a theory, and the rest still has to be rebuilt in each head. Naur’s answer to how that happens is “close contact”: the new programmer works with “the programmers who already possess the theory” (Naur 1985). Mathematics has long-standing ways of doing this: seminars, textbooks, apprenticeship. What would the software equivalent be for the output of an agent, where there is no one who possesses the theory to be in close contact with?
In science
Here I think science gives me a bigger leverage. In Where do correctness and understanding live in research software? I argued that for scientific software, the theory of the program is, in large part, the theory of physics. So in scientific computing the theory of the program closely follows our scientific understanding, and we have been laying that out for a long time.
Take the chain from that post: the model equations, the approximation, the algorithm, the code. The equations are in the paper. Why this approximation, and where it breaks down, is in the numerical analysis. Tapio Schneider describes trusted predictions as the output of “an auditable chain […] whose links have known limits of validity and can be tested individually.” (Schneider 2026) (see my link post). Each link of that chain can be written down, so what is left to be rebuilt in a head is much smaller than for, say, Word. Literate programming would put it right next to the code (Knuth 1984).
So for science, I’d claim the theory can mostly be laid out.
Can it be rewarded?
As I argued in the dead programs post, correctness has verifiable signals (tests, specifications, static analysis) that reinforcement learning with verifiable rewards (RLVR) and agents in a loop can act on. Understanding doesn’t, and Naur says “the very notion of qualities such as simplicity and good structure can only be understood in terms of the theory of the program” (Naur 1985), so complexity metrics are at best proxies.
The only signal I know of that isn’t a proxy is Naur’s own test. A program is dead when “demands for modifications of the program cannot be intelligently answered.” (Naur 1985) So whether an output lays out a theory shows up in whether the next modification goes well. That’s a delayed reward, measured on future changes, and as far as I know that’s what reinforcement learning is worst at.
There are existing attempts at delayed signals:
- SlopCodeBench has agents repeatedly extend their own code as the specification evolves over checkpoints, and measures verbosity and structural erosion (Orlanski et al. 2026). I think this measures decay rather than theory.
- In mathematics, reuse: a definition in
mathlibthat simplifies everything built on it, as I argued in What should mathematics reward?. - Comprehension checks on the humans. In Anthropic’s randomized study of 52 mostly junior engineers learning a new Python library, the group using AI averaged 50% on a quiz afterwards, against 67% for the group coding by hand (Shen and Tamkin 2026). It is indirect: it measures the human’s understanding, not whether the output lays one out. But if the LLM had clearly laid out the theory to the humans, why would they score lower? It is a hint that, at least, it doesn’t build and transmit an understanding.
Two camps
Some mathematicians claim it can’t be done. The Fields medallists’ declaration says “solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight”, and that “without the willing mathematicians who must take care of their development and integration into the mathematical canon, AI-conceived ideas would never become fully alive” (Avila et al. 2026). That sounds a lot like Naur’s dead program. Bryna Kra writes that machine-generated mathematics “does not replace the role of human expertise”, and calls for rewarding “what we truly value, and not just what we can easily measure” (Kra 2026). That is right, though I think it takes for granted that what we truly value, understanding included, can’t be measured. Schneider, from climate science, has “humans, for now, remaining essential for extracting understanding” (Schneider 2026). None of them claims to know ML. Their position is that mathematicians need to emphasize understanding because that is where only a human can do it, and I think this implicitly assumes ML can’t solve that problem. Henry Cohn is more careful: “Mathematical truth may be easier to train for than effective communication with humans” (Cohn 2026).
As for ML, I have heard hopes that it can be done. But I have never heard a constructive reason that it can.
On the other hand, it might be a descriptive problem. John von Neumann, in a talk on computers in 1948, answered the canonical question from the audience, whether a mere machine can really think: “You insist that there is something a machine cannot do. If you will tell me precisely what it is that a machine cannot do, then I can always make a machine which will do just that!” E. T. Jaynes, who was in the audience, put it as “the only real limitations on making ‘machines which think’ are our own limitations in not knowing exactly what ‘thinking’ consists of.” (Jaynes 2003)
So when one says something cannot be done in AI, it may just be that we don’t know how to write down the criteria, in a sense that we ourselves don’t really understand “it” yet, whatever “it” is. In this case “it” is understanding itself.
So back to the first question. I claimed a theory can be laid out in science, and mathematicians claim that understanding is exactly what a machine can’t do in mathematics. Science is close to mathematics, so I’d expect the two answers to be similar, whichever way it goes.1
References
Footnotes
There is a caveat. Which field requires more understanding? Which field’s understanding is harder? And is it even the same thing? Schneider separates episteme, “roughly explanatory understanding”, from techne, “roughly the craft of predicting and making” (Schneider 2026), and when I read his post, understanding in science came out as a model that compresses and generalizes better, which is not obviously what mathematicians mean by it.↩︎