AI Affairs, home

Thursday 8 October 2026

Technology

OpenAI releases 722 mathematical manuscripts from an internal model

The collection groups the papers into 372 families and includes computer-checkable versions of many proofs. OpenAI estimates that an average result used computing equivalent to roughly three hours of ChatGPT Pro thinking.

OpenAI’s headquarters at 1515 Third Street beneath crisscrossing overhead wires
Photo: Coolcaesar, CC BY 4.0, via Wikimedia Commons (cropped)

OpenAI released a collection of mathematical results produced by an internal frontier model on 6 October 2026, publishing manuscripts and supporting material in a public GitHub repository. The collection contains 722 manuscripts arranged in 372 families, Unite.AI reported. Alongside the papers, OpenAI has provided computer-checkable versions of many proofs, examples of the model’s reasoning and estimates of the computing used to obtain the results.

Key points

  • The public collection contains 722 manuscripts grouped into 372 families.
  • OpenAI estimates that an average result used computing equivalent to roughly three hours of ChatGPT Pro thinking.
  • Many proofs have versions written in Lean so a computer can check them.
  • An independent advisory group has recommended fuller disclosure and scholarly repositories for AI-generated mathematical results.

The 372 families in OpenAI’s collection

The manuscripts grew out of evaluations on open research problems during the development of an unreleased model. A family can bring together a main result with related arguments, consequences or alternative proofs, and the catalogue sorts families by mathematical discipline. The repository carries PDFs and source files as well as instructions for building and citing individual manuscripts.

Unite.AI described the results as being at different stages of verification. Some papers have accompanying formal proofs and others do not. Its account of the README says some results without a formal version could have problems, and that corrections will be recorded as revisions while earlier public versions remain available. That matters for a collection intended to be read and cited: a manuscript can circulate while its argument is still being examined.

OpenAI says it has established procedures for revisions and citations in the GitHub repository. Those procedures give the papers a public history as mathematical arguments are checked, corrected or built upon; they do not put every paper at the same stage of verification.

Lean proofs and ChatGPT Pro compute

OpenAI has included formalizations of many proofs in Lean. A mathematical proof written in ordinary language asks a reader to follow its reasoning, whereas a Lean formalization expresses the argument in a form a computer can check. Think of the difference between reading a set of directions and having every turn checked against a map: the check depends on the directions having been written precisely enough. OpenAI’s collection includes that second form for many proofs, though the README cautions about possible issues among results without it.

Checking an answer to a maths problem could involve asking whether its proof has a computer-checkable version, rather than relying on how convincing the explanation sounds. That would depend on such a version being available; OpenAI says many of these proofs have one.

The model was given approximately 4,000 problems during the evaluation, according to the repository account reported by Unite.AI. OpenAI says the average published result used computing equivalent to roughly three hours of ChatGPT Pro thinking. The estimate describes computing use in equivalent hours of ChatGPT Pro thinking, not a monetary price. The repository also includes statistics on attempted problems and 10 abridged summaries of the model’s reasoning.

According to the repository account, most results followed the same evaluation procedure, though work on a region associated with the Riemann zeta function and a proof concerning CM abelian varieties were exceptions. A manuscript about the zeta function was edited by a human for readability. The papers and their supporting files therefore give people building on the results material to inspect, while the described exceptions matter when comparing how individual papers were produced.

The internal model and the September 8 proof

The model remains internal. On 8 September, OpenAI said its internal system had produced a proof, formalised in Lean, that initially smooth fluid flow in three dimensions can reach a singularity in finite time under the Navier–Stokes equations. OpenAI described that result as resolving one of the Clay Mathematics Institute’s Millennium Prize Problems and said the model was significantly more capable than GPT-6 Astra.

That earlier announcement gives the present release a research context: OpenAI is putting a much larger body of mathematical writing into public view while the model that produced it remains unreleased. The repository includes manuscripts, formalizations and accounts of how results were obtained, rather than access to the system itself.

The advisory group’s 29 September recommendations

OpenAI says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study about how to share the work. The group published recommendations on 29 September after receiving more than 600 replies from the mathematical community, Unite.AI reported. It urged labs to put results in scholarly repositories outside their control, assign persistent identifiers and document modifications appropriately.

The group also recommended publishing the model name, prompts, a summary of the reasoning, the time taken and an estimate of computation costs for each result. For large releases, it called for an account of how problems were selected and how many comparable problems the model tried without solving. OpenAI’s GitHub release includes compute estimates and attempted-problem statistics, while the company says it is exploring community-hosted alternatives that meet the group’s guidelines.

The advisory group recommended support for work that helps mathematicians understand AI-generated results, including workshops and longer-term working groups. It said existing nonprofit institutions should decide which efforts receive that support, rather than the labs producing the results.

Sources

Topics: Foundation models, Inference