What does understanding mean in astronomy?

Link
LLM
Astronomy
Research ethics
Research software engineering
Jess Werk, chair of Astronomy at the University of Washington, on protecting the Ph.D. from AI — and how differently astronomy, CMB science and RSE might mean “understanding”.
Author

Kolen Cheung

Published

October 3rd, 2026

As scientists, we design our experiments around an unattainable ideal: an objective and mindless observer from nowhere, studying a reality that exists separately from our subjective understanding of it. As creatures who study the cosmos from a relatively tiny rock, […] we astronomers appreciate the physical impossibility of the observer from nowhere (and yet we endeavor to achieve it!). […] Intelligibility is the biological mind’s responsibility; every experiment we design must be understood by someone because that is what gives science its purpose. Generative AI attempts to embody the mindless observer from nowhere. Its achievements in both mathematics and science reveal the incompleteness of this long-held ideal and underscore the importance of human scientists and mathematicians, whose minds make results and proofs meaningful.

Frankly, I don’t fully understand what the “unattainable ideal” really means. Unpacking it: the observer from nowhere is Thomas Nagel’s view from nowhere, an observation that doesn’t depend on who makes it or from where, and astronomers literally can’t have one, observing from one rock inside one galaxy. The way I understand it, the objective of science is to understand nature (the objective thing here), which is interesting in itself: it almost means science itself is just a “hobby” of intelligent beings. We don’t have to understand it, but we want to. And because of that, it connects intelligibility to biological beings.

That raises more questions than it gives clarity. Some would argue that in science they don’t care about understanding but about prediction, because prediction has value: it is not just an intellectual hobby, it has utility. That is the distinction drawn in episteme and techne.

In astronomy, the result plays the role of the theorem: a paper is judged on what it found, how significant it was, and whether it found it first, far more than on what the finding means. The result serves as a proxy for understanding, and that proxy is now a problem (Kra, B., 2026). Across academia, incentive structures have long rewarded productivity over understanding, and generative AI now makes productivity easy to manufacture.

Werk seems to be saying that the proxy problem Kra describes also exists in astronomy. But that doesn’t sound right: making such a discovery in astronomy is so expensive that pure AI won’t solve it. Expensive experiments are protected exactly because the barrier to entry is high, unlike math and software. That’s why software eats the world.

At our Astronomy faculty retreat on September 17, we discussed a scenario modeled on what is happening in mathematics. Briefly, OpenAI posts a preprint, press release, and public decision log of 40,000 steps reporting a five-sigma detection of an evolving dark energy equation of state, inconsistent with a cosmological constant, with error bars a factor of 2.5 tighter than anything our community could achieve from the same public data, drawn from a future survey in which our department has already invested heavily. None of the faculty were especially fazed and none thought the scenario to be implausible.

OK, this starts to make sense. The scenario has no barrier to entry from building any hardware to make observations. It takes existing publicly released data, builds a new model (or cherry-picks existing ones) with a new analysis that nobody else has thought of or had the time to do, and lands on a new discovery.

I’m skeptical, primarily because of the cost of a parallel search, given the amount of “tool use” here: each idea requires a full end-to-end analysis (OK, maybe some parts can be reused). There are two possibilities: an AI lab launching such a sweep, or an individual or small group using a tool from an AI lab to launch it.

An AI lab, while it has the resources at its disposal, must be doing it to demonstrate capability, so it must be something much harder for an individual to do. Think Navier–Stokes level. So the approach would likely be similar: a massive parallel sweep to try many ideas, which is what an individual can’t afford. Then not only do they need to spend the equivalent of tens of millions of dollars’ worth of compute on the intelligence alone, but they also need to spend a huge amount of compute per idea: for the sake of argument, say the equivalent of a million NERSC hours to process petabyte-scale data. That would be orders of magnitude more than they needed to spend on Navier–Stokes.1

For an individual, the idea is the same, just at a smaller scale. The bottleneck is still getting an allocation of compute granted, and the grant proposal should have something concrete to write in the first place. Unless compute is abundant, you can’t just let an AI try out ideas and stumble upon an answer, as in the case of Navier–Stokes or similar achievements. And compute is not getting more abundant: AI demand is already driving up the cost of RAM (TrendForce expected conventional DRAM contract prices to rise 90–95% in the first quarter of 2026) and starting to on CPUs as well (Intel’s server processors in China up more than 10% by February 2026, with agentic AI named as a driver).

A second issue with this idea is p-hacking: if you try out 10,000 ideas and get a 5σ discovery, its p-value is not really that.2

The scenario’s public decision log of 40,000 steps would let you count the tries, if it includes all of them. But one would likely, mistakenly, only show the thoughts of the successful run, like how the Navier–Stokes proof was reported: the other almost-but-not-quite proofs never became papers saying why not that idea.

But go on.

[…] A paper is rarely dedicated to showing how a particular idea turned out to be wrong after a careful analysis. “Nobody has time for that,” we think. […] A cognitive sanctuary, therefore, must reward the effort and celebrate a process that includes failure. It is a space that encourages unhurried thinking with plenty of room for rabbit holes and the occasional mad hatter.

Interesting: on top of digesting and understanding, which the mathematicians advocate, this includes failure, which actually makes sense for a science. In science, at least in principle, making falsifiable claims is more important than making “correct” ones. So Werk is pointing out one problem of our reward system: it doesn’t incentivize those kinds of activities. As far as I know, no mathematician is saying: show me your almost-proofs that don’t work too, so that we understand more about what wouldn’t.

Unlike a mad hatter, agentic AI creates a chain of probabilistic decisions and steers users away from the improbable. […] Practice builds a tacit element of understanding that my field calls physical intuition. Nothing yet shows that physical intuition can be developed through closed-loop agentic AI conversations, and recent evidence points in the opposite direction (Bastani et al. 2025). Interrogation is best practiced among other scientists who can offer surprising alternatives and explanations (sometimes incorrect!), while an AI agent’s alternatives are drawn from the distribution of what has already been written.

And if the point of science is to explore, an agentic AI that is not good at creativity, at going out of distribution, reduces the chance of sampling the improbable. Sort of like a data scientist forgetting to do importance sampling.

This seems to parallel AI not being able to generate mathematical understanding. Here it’s physical intuition.

Academic departments bring together kindred spirits and support many of the structures a cognitive sanctuary needs: we host seminars, discussions, hack-a-thons, and community events designed to honor thinking minds.

Werk then gives practical suggestions in three groups, Ph.D. processes, faculty reward systems and department community, that I can only describe as cultivating understanding: assess the student live, at the board and without AI, rather than the text they hand in; stop pushing students to publish before the interpretation has been worked out; and give faculty credit for mentoring and for showing up. I think it is not too dissimilar to how mathematicians are advocating understanding, and Werk notes that the Harvard Summit on PhD Math Education in the Age of AI “reached many of the same conclusions”. The Fields medallists’ declaration has the most succinct list of what the math community does about it:

The most precious resources of our profession are students and ideas, and these we nurture with great care. […] For students we often suggest problems with the core intention of developing skills making them well-positioned for advances in research and elsewhere. Our ideas we disseminate in talks, private discussions and careful writeups, connecting them to the previous ideas of others. These processes invariably take time and are based on human interaction.

In an optimistic future where AI-driven discoveries and innovations must be understood and tested, […] Ph.D. students find joy as their engaged human professors challenge them over and over again until what they know expands, and they can vouch for all of it.

Werk and I must have a very different idea of how “understanding” is acquired in astronomy. For example, I heard someone make the argument that astrophysicists listening to an astronomer presenting their discovery would be puzzled that they got into so much detail about their observations, and how strange and unique that was, and when someone asked something to the effect of “what does it really mean physically?”, communication between the two crowds seemed to break down. It doesn’t translate.

Also, from my own experience as a (former) CMB scientist, we got so far into the details of our scientific methods, including blind analysis and the detection, estimation and mitigation of systematic biases, and finally landed on a measurement, with statistical and systematic uncertainties. Where is the “understanding”, though? Are we saying we have an understanding of the instrumentation techniques, of the systematics, of the data analysis methods and their limitations, and of the implications of the measurement? And in the rare case that a measurement leads to a discovery (say, confirming that primordial gravitational waves exist, or even starting to falsify inflationary models), where does the understanding originate? With the theorists who proposed these models long ago, or with this particular experiment? What is the understanding? That among models A, B and C, I now know B is the right one? That sounds more like an answer to me, where the understanding of what B means came long ago.

One dimension, for example, is whether the student can state the competing interpretations of the result and say what measurement would distinguish them.

The “what would” part is indeed intellectual and creative: it designs a future experiment. But I’m thinking more of my own field, measuring primordial B-modes, where what you are measuring has become standardized. Primordial B-modes have been searched for many years, still without a detection, so it is a well-worn path: someone has “long predicted” a bunch of models, and you are falsifying against them. One of Planck’s papers does that kind of analysis (Planck Collaboration 2020). So if the next experiment to build, in the case of CMB, is just better sensitivity to beat the noise, is that real understanding? You may say doing so is very hard, and still requires a lot of creativity and understanding of the instrumentation to get right. Then it goes back to the point that understanding seems to mean very different things here.

Now, all this is not a criticism. Probably I’m just thinking out loud, and the conclusion seems to be that even what “understanding” means is very different in different fields (math vs. physics vs. astronomy vs. astrophysics), even in different sub-fields (e.g. CMB), or with different expertise intersecting with a sub-field (say theorists, data scientists, instrumentalists, etc.).

Thank you to Professor Xavier Prochaska, my postdoc mentor who frequently had me go to his chalkboard with my ideas. He came to me over a year ago with a claim that, under his guidance, Claude could write a Ph.D. thesis in a month that would be better than that of an average Ph.D. student in Astronomy. I did not believe him at the time. I do now, but I have decided that the thesis is not the point of the Ph.D. The student is.

They can say that. But translate almost the same kind of intelligence and capability to RSE: “Under a PI’s guidance, Claude could write research software in a month that would be better than that of an average RSE.” Then the reason for raising a student is invalid here, basically because theirs is vertical: today’s student becomes tomorrow’s professor. Ours is horizontal: the PI reaches out horizontally to collaborate with external expertise.

That probably goes back to my earlier question of what our value is, where I proposed communication and translation, without which an AI cannot write research software better than we do with AI. We are not arguing against using AI, but about who is using it: if I use it better than you do because I know how to extract the user stories and turn them into software design criteria, that defines our value. It’s why you would hire an RSE rather than buy a Claude Code subscription. RSE is in a vulnerable position that astronomy and mathematics, with their vertical line from student to professor, don’t share, because this is not clearly communicated to all potential PIs (which was true even pre-LLM).

References

Planck Collaboration. 2020. “Planck 2018 Results. X. Constraints on Inflation.” Astronomy & Astrophysics 641: A10. https://doi.org/10.1051/0004-6361/201833887.

Footnotes

  1. For scale: taking a NERSC hour as roughly a CPU core-hour, which costs a few cents in the cloud, a million of them is tens of thousands of dollars per idea, so 10,000 ideas would be hundreds of millions of dollars. OpenAI’s output tokens for Navier–Stokes were worth about $6.5 million at public API prices.↩︎

  2. Physicists call this the look-elsewhere effect. The simplest correction is Bonferroni’s, which multiplies the p-value by the number of tries; Holm–Bonferroni is the step-down version for testing many hypotheses at once, and is the same for the most significant one. A one-sided 5σ is p ≈ 2.9 × 10⁻⁷, so after 10,000 tries it is p ≈ 2.9 × 10⁻³, about 2.8σ. To claim 5σ, the best of the 10,000 would have to reach about 6.6σ on its own. Tries that reuse most of an analysis are not independent, so Bonferroni overcorrects, but not by orders of magnitude.↩︎