What do research software engineers uphold?

Talking about AI in the RSE community, borrowing from mathematics

Dr. Kolen Cheung, Research Software Engineer

Research Software & Analytics Group, University of Exeter

October 5th, 2026

Who I am

  • Research software engineer, Research Software & Analytics Group, University of Exeter.
  • PhD in computational cosmology.
  • A mathematician at heart.
  • Maintainer, co-maintainer and contributor of open source projects, including in the pandoc ecosystem.
  • I explore and apply agentic engineering in my work, and write about AI literacy, safety and ethics.

What I do

  • At the intersection of science, statistics, software engineering, and lately agentic engineering.
  • How computational and statistical techniques help us answer questions in science: make the infeasible feasible, and make sure it is correct along the way.
  • My blog, blog.kolen.dev: reproducibility, JIT compilers, statistical methods, agentic engineering, and lately how AI is impacting mathematics, and the parallels with RSE as a profession.
  • An open source advocate and practitioner. Agents now help me maintain projects I’d otherwise not have the time for.

My observations on reactions to AI

  • Are you “pro-AI” or “anti-AI”?
  • Multidimensional: usefulness · understanding · sustainability · safety · responsible use · mental health · jobs
    • Each can lead to “use” or “don’t use”. Two people who both say “don’t” may disagree on why.
  • Collapsing it to sides makes the conversation hard. Some are angry at people who use AI. People who use it stop talking about it.
  • The vicious cycle: hostility, users go quiet, less shared practice, deeper divide.

This room likely contains grief, enthusiasm, anger, denial and exhaustion.

— Arfon Smith, RSECon26 keynote, Where does the rigour go? (2026)

Between two neighbours

Software engineering

  • Years ahead with coding agents, and the most direct impact on us.
  • Dark factory:
    • Correctness: Shipping code nobody has read is a “when” and not an “if” (Charity Majors).
    • Understanding: Those on call are “seeing melting mental models” (Orosz 2026).

Mathematics

  • Hit hardest and fastest: harder and harder problems solved this year, up to a machine-checked Navier–Stokes blowup in September (OpenAI 2026).
  • The Fields medallists (2026): “solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight”.
  • RSE shares characteristics between the two.

I know mathematics and physics best, so my examples come from scientific computing. RSE outside STEM exists too. See Where do correctness and understanding live in research software?

Must we read every line?

  • Some RSEs hold that you must read every line an agent writes.
  • Nobody audits every line of assembly a compiler emits. We check the source, the tests and the result.
  • Mathematicians don’t read a Lean proof either. The kernel checks it. They check the statement, any sorry left, etc.
  • My claim: for agent-written code, care about correctness, not the exact implementation. But software abstractions leak, and we have no kernel, so the art is channelling where a human looks.
  • Agree or not, this is what we need to talk about before we can say what we value.

“You never delete human verification. You relocate it somewhere smaller: assembly → source → spec → a harness you can audit.” (JIT talk, August 2026)

Borrow from mathematics: values

  • Within days of Navier–Stokes, the Fields medallists said what they value (2026):
    • understanding · exposition · credit · “The most precious resources of our profession are students and ideas”
  • They aren’t conceding, because they never moved the goalpost: “The product of mathematics is clarity and understanding” (Thurston 2010).
  • RSE is young (2012), and sits between neighbours that pull different ways. We haven’t said what we value.
  • If our value is efficiency, we lose at some point, then move the goalpost.
  • My candidate: what understanding is to mathematics, communication is to RSE. User stories → specifications → implementation → back to the researchers.

This is my candidate, as a question. The plan is for the community to answer it.

The plan: listen, measure, discuss, shape

Listen

~25 conversations with people who hold strong views, across the country and across dimensions. Anonymised in what I publish.

Measure

A community survey: what we think we should do, what we actually do, and what we tell others.

Discuss

Blog, podcast, talks, workshops. Including how AI affects our wellbeing, on every side.

Shape

The community says what RSEs value. A statement only if one emerges.

Starting from what exists: declarations, papers and blogs on AI in the RSE community.

Feasibility

  • Time: 0.1 FTE over 15 months (about 240 hours).
  • Tentative targets: 25+ conversations, one survey, a published synthesis, 4–6 talks or podcasts.
  • What the Fellowship adds: reach for the survey and conversations, a cohort, a mentor.

What we have at the end

  • People on every side who have been heard, and data on where we actually stand.
  • A shared answer, or a clearer disagreement, on what RSEs uphold.