What do research software engineers uphold?
Talking about AI in the RSE community, borrowing from mathematics
Dr. Kolen Cheung, Research Software Engineer
Research Software & Analytics Group, University of Exeter
October 5th, 2026
This is my application to the Software Sustainability Institute Fellowship 2027: a six-minute screencast of the slides above. Below each slide is what I say over it.
Who I am
- Research software engineer, Research Software & Analytics Group, University of Exeter.
- PhD in computational cosmology.
- A mathematician at heart.
- Maintainer, co-maintainer and contributor of open source projects, including in the pandoc ecosystem.
- I explore and apply agentic engineering in my work, and write about AI literacy, safety and ethics.
My name is Kolen Cheung. I’m a research software engineer at the University of Exeter. My PhD is in computational cosmology, and I’m a mathematician at heart. I maintain and contribute to open source projects, including in the pandoc ecosystem. I use AI agents in almost all of my work now, so I see both what they can do and what they cost.
What I do
- At the intersection of science, statistics, software engineering, and lately agentic engineering.
- How computational and statistical techniques help us answer questions in science: make the infeasible feasible, and make sure it is correct along the way.
- My blog, blog.kolen.dev: reproducibility, JIT compilers, statistical methods, agentic engineering, and lately how AI is impacting mathematics, and the parallels with RSE as a profession.
- An open source advocate and practitioner. Agents now help me maintain projects I’d otherwise not have the time for.
What I care about is how computational and statistical techniques help us answer questions in science: making the infeasible feasible, and making sure it is correct along the way. I write about this on my blog: reproducibility, JIT compilers, statistical methods, agentic engineering, and lately, what AI is doing to mathematics, and what that means for us as RSEs.
My observations on reactions to AI
- Are you “pro-AI” or “anti-AI”?
- Multidimensional: usefulness · understanding · sustainability · safety · responsible use · mental health · jobs
- Each can lead to “use” or “don’t use”. Two people who both say “don’t” may disagree on why.
- Collapsing it to sides makes the conversation hard. Some are angry at people who use AI. People who use it stop talking about it.
- The vicious cycle: hostility, users go quiet, less shared practice, deeper divide.
This room likely contains grief, enthusiasm, anger, denial and exhaustion.
— Arfon Smith, RSECon26 keynote, Where does the rigour go? (2026)
Are you pro-AI or anti-AI? I reject the question. There are many dimensions, from usefulness and understanding to mental health and jobs. Each can lead to use or don’t use, and two people who both say don’t may disagree on why.
Two sides didn’t help US politics or the COVID-19 response. Some are angry at people who use AI, so those who use it stop talking about it, and the divide deepens. As Arfon Smith said at RSECon26 (2026), this room likely contains grief, enthusiasm, anger, denial and exhaustion.
Between two neighbours
Software engineering
- Years ahead with coding agents, and the most direct impact on us.
- Dark factory:
- Correctness: Shipping code nobody has read is a “when” and not an “if” (Charity Majors).
- Understanding: Those on call are “seeing melting mental models” (Orosz 2026).
Mathematics
- Hit hardest and fastest: harder and harder problems solved this year, up to a machine-checked Navier–Stokes blowup in September (OpenAI 2026).
- The Fields medallists (2026): “solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight”.
- RSE shares characteristics between the two.
I know mathematics and physics best, so my examples come from scientific computing. RSE outside STEM exists too. See Where do correctness and understanding live in research software?
RSE sits between two neighbours. Software engineering is years ahead with coding agents. There, Charity Majors says shipping code nobody has read is a “when” and not an “if”. She also says those on call are “seeing melting mental models”.
Mathematics was hit hardest and fastest. This year AI solved harder and harder problems, up to last month, when OpenAI announced a machine-checked blowup for the Navier–Stokes equations, a Millennium Prize Problem. The Fields medallists answered that solving problems is only a tool and proxy for understanding.
So one neighbour says ship it unread, the other says understanding is the product, and we are pulled both ways.
Must we read every line?
- Some RSEs hold that you must read every line an agent writes.
- Nobody audits every line of assembly a compiler emits. We check the source, the tests and the result.
- Mathematicians don’t read a Lean proof either. The kernel checks it. They check the statement, any
sorryleft, etc. - My claim: for agent-written code, care about correctness, not the exact implementation. But software abstractions leak, and we have no kernel, so the art is channelling where a human looks.
- Agree or not, this is what we need to talk about before we can say what we value.
“You never delete human verification. You relocate it somewhere smaller: assembly → source → spec → a harness you can audit.” (JIT talk, August 2026)
Here is a claim some of you will disagree with. Some RSEs hold that you must read every line an agent writes. But nobody audits every line of assembly a compiler emits. Mathematicians don’t read a Lean proof either. The kernel checks it, and they check the statement.
So I think for code written by an agent, we should care about correctness, not the exact implementation. Software abstractions leak, and we have no kernel, so the art is channelling where a human looks. Agree or not, this is what we need to talk about before we can say what we value.
Borrow from mathematics: values
- Within days of Navier–Stokes, the Fields medallists said what they value (2026):
- understanding · exposition · credit · “The most precious resources of our profession are students and ideas”
- They aren’t conceding, because they never moved the goalpost: “The product of mathematics is clarity and understanding” (Thurston 2010).
- RSE is young (2012), and sits between neighbours that pull different ways. We haven’t said what we value.
- If our value is efficiency, we lose at some point, then move the goalpost.
- My candidate: what understanding is to mathematics, communication is to RSE. User stories → specifications → implementation → back to the researchers.
This is my candidate, as a question. The plan is for the community to answer it.
What I’d borrow from mathematics is values. Within days of Navier–Stokes, the Fields medallists said what they value: understanding, exposition, credit, and their students. They never moved the goalpost: Thurston wrote, long before AI, that “the product of mathematics is clarity and understanding”.
RSE is young, the term is from 2012, and I don’t think we have said what we value as coherently as the mathematicians. If our value is efficiency, we lose at some point, and then we have to move the goalpost.
What understanding is to mathematics, I think communication is to RSE as a profession. We extract user stories from our researchers into concrete specifications and requirements. We implement them, which is the problem-solving part where coding agents could replace us. And we communicate back how to use it and adapt it to their research. That said, these are questions for the community, through which we crystallize our values.
The plan: listen, measure, discuss, shape
Listen
~25 conversations with people who hold strong views, across the country and across dimensions. Anonymised in what I publish.
Measure
A community survey: what we think we should do, what we actually do, and what we tell others.
Discuss
Blog, podcast, talks, workshops. Including how AI affects our wellbeing, on every side.
Shape
The community says what RSEs value. A statement only if one emerges.
Starting from what exists: declarations, papers and blogs on AI in the RSE community.
So my plan for the Fellowship has four steps: listen, measure, discuss and shape.
First, listen. Around 25 conversations with people who hold strong views, across the country and across these dimensions. I’d have each conversation with a counsellor’s hat on, and anonymise them in anything I publish.
Then measure. From those conversations, I’d design a community survey to measure where we stand across these dimensions. For example, should we audit LLM output line by line? Do you? Do you tell people when you don’t?
Then discuss, on my blog, in podcasts, talks and workshops, including how AI affects our wellbeing, on every side. This is also where my own view goes, and I’ll say when I’m advocating.
Finally, shape: the community says what RSEs value. If a statement emerges with consensus, I’d do it the Fields medallists’ way: a few invited signatories, and many endorsers.
Feasibility
- Time: 0.1 FTE over 15 months (about 240 hours).
- Tentative targets: 25+ conversations, one survey, a published synthesis, 4–6 talks or podcasts.
- What the Fellowship adds: reach for the survey and conversations, a cohort, a mentor.
I can give this 10% of my time over the 15 months, about 240 hours, from the time our group sets aside for team initiatives and personal development. Tentatively: 25 or more conversations, one survey, a published synthesis, and four to six talks or podcasts.
What the Fellowship adds is reach: conversations across the community, and a survey that goes as wide as possible. It also brings a cohort and a mentor, who would be invaluable to brainstorm approaches and illuminate blind spots.
What we have at the end
- People on every side who have been heard, and data on where we actually stand.
- A shared answer, or a clearer disagreement, on what RSEs uphold.
At the end, I hope people on every side have been heard, and we have data on where we actually stand. And we have a shared answer on what RSEs uphold, or at least a clearer disagreement. Thank you.