AI Affairs, home

Saturday 10 October 2026

Technology

Mathematicians find two gaps between OpenAI’s Navier-Stokes proof and its code

In one passage, the machine-checked code proves a weaker estimate than the written argument states. The mathematicians do not claim the written proof is wrong.

The Pioneer Building in San Francisco behind trees and parked cars
Photo: HaeB, CC BY-SA 4.0, via Wikimedia Commons (cropped)

OpenAI’s Navier-Stokes result, announced on 8 September, is under scrutiny after three mathematicians identified two places where its written proof diverges from the version checked in Lean, Yellow reported. In one place, the code establishes a weaker estimate than the paper states. Alexander Bastounis of King’s College London and Fabian Circelli and Anders Hansen of the University of Cambridge documented the differences without claiming that the written proof is incorrect.

Key points

  • Three mathematicians found two divergences between OpenAI’s written argument and its Lean code.
  • One coded estimate is weaker than the corresponding statement in the paper.
  • OpenAI’s manuscript catalogue fell from 722 to 719 after it withdrew three papers over a sign error.
  • The mathematicians say a computer check cannot take the place of human peer review.

Two differences in OpenAI’s Lean proof

OpenAI produced a proof written for mathematicians and a formalised version that a computer could check. The distinction matters because Lean checks the statements entered into it, not whether those statements faithfully reproduce every passage of a separate paper. A checked proof is rather like a chain tested link by link: the test concerns the chain in front of the examiner, not a description attached to it.

Bastounis, Circelli and Hansen found that the two versions do not match at two points. The reported difference in an estimate is specific: the Lean code proves a weaker statement than the written paper makes at that point. The authors say Lean confirms the final theorem in the formalised version, but that this does not validate every intermediate step a mathematician encounters in the prose.

Checking a calculation before including it in a written report would still involve comparing the explanation with what the computer has verified. A passed check could support the coded steps while leaving a stronger statement in the explanation to be assessed separately.

The mathematicians’ objection concerns the relationship between the two proofs, rather than a demonstrated error in the final written result. Their conclusion is that machine checking cannot replace peer review when a result is presented in prose as well as code. That is a particular demand on readers of AI-generated mathematics: they must assess the argument being offered to them, even when a related formal argument passes a check.

Hansen told New Scientist on 8 October that large language model-generated proofs would have to be read by people, creating an “enormous extra burden on mathematicians”. His concern applies beyond this one estimate if models can produce written results faster than mathematicians can examine the correspondence between their claims and their supporting arguments.

OpenAI’s catalogue falls to 719 manuscripts

The Navier-Stokes examination arrives alongside OpenAI’s larger release of mathematical work. The company published 722 manuscripts on 6 October, as AI Affairs reported when the collection appeared. OpenAI subsequently withdrew three papers, taking the catalogue to 719, after a sign error undermined one argument and two papers dependent on it, Yellow reported.

OpenAI’s repository describes the manuscripts as results from testing an unreleased internal model on open research problems. According to Yellow, summaries of the model’s reasoning accompany 10 results, while about 42 per cent of the collection’s top-line results have been formalised in Lean.

The withdrawn papers present a separate problem from the mismatch in the Navier-Stokes work. A sign error broke an argument and affected two others that relied on it; the three mathematicians, by contrast, identified differences between a written argument and its coded counterpart. In both cases, examining the material beyond a paper’s stated result is part of determining what the released work supports.

The 29 September mathematics guidelines

A nine-member Advisory Group on Mathematics and Artificial Intelligence issued guidelines on 29 September asking laboratories to stop testing advanced mathematical problems on proprietary models. OpenAI’s description of the unreleased model used for its manuscript collection puts that request in direct contact with the way the work was produced. The advisory group has said that assessing OpenAI’s adherence to its guidance is a task for mathematicians at large.

Terence Tao wrote on 6 October that some problems were being solved by “AI prompters” without sufficient grasp of the results to field questions or deliver talks. The concern is distinct from whether a machine can verify a formal statement. It concerns who can explain and defend a result after it has been circulated for other mathematicians to use.

The scope of OpenAI’s Navier-Stokes claim is also contested. The equations describe the motion of fluids such as water and air. OpenAI says it proved that a singularity can arise under a particular external force, while mathematicians argue that the central question is whether one can occur without such a force, Seoul Economic Daily reported on 9 October. A singularity is a point at which the mathematical description produces an abnormal value, such as a velocity rising without bound.

Twenty-five Fields Medal winners issued a declaration on 11 September warning that mass-produced AI results could harm mathematics. OpenAI has established an independent advisory committee and says it will seek ways to use AI to support mathematical understanding. It has refused the request to stop testing advanced problems with its own models, saying that work is critical to developing tools for the field, Seoul Economic Daily reported.

Sources

Topics: Foundation models, Safety