On October 6, 2026, OpenAI launched a public repository named openai/math on GitHub, releasing 722 mathematics manuscripts under the Apache-2.0 license, grouped into 372 problem families, accompanied by Lean formal proof files and model reasoning summaries. The results were produced by an undisclosed internal frontier model, which, according to unite.ai, is significantly stronger than GPT-6 Astra. This is the largest public collection of AI mathematical output to date.
The production method for this batch of manuscripts was relatively fixed: the model was fed about 4,000 open mathematics problems, and through significance screening and grouping of results, 722 manuscripts were ultimately produced. OpenAI disclosed that the compute consumed by each selected result on average is equivalent to about three hours of ChatGPT Pro thinking.
Four Core Result Types, of Uneven Value
According to technical reviews from Chinese media outlets such as NetEase Tech, four categories in this batch of manuscripts have drawn the most attention from the mathematics community:
- The Quasi-Riemann Hypothesis: The manuscript claims to prove that the Riemann zeta function has no zeros in the region where the real part is greater than 7/8. After seeing this result, Rutgers University mathematics professor Alex Kontorovich publicly said: "If a human had done this, it would be an immediate Fields Medal, no question." The full Riemann Hypothesis is not thereby declared solved, but this extension of a fixed-width zero-free region is considered technically far beyond previous expectations. The result is accompanied by Lean formalization materials.
- A Special Case of the Hodge Conjecture: A proof is given for the rational Hodge conjecture on CM abelian varieties. The Hodge conjecture is one of the seven Millennium Prize Problems; this addresses a special but important class, and no complete Lean formalization has been provided.
- The Unique Games Conjecture: The manuscript claims to give a proof and derives approximation hardness bounds for problems such as Max-Cut. Under the assumption that P≠NP, this means that the approximation guarantees of efficient algorithms for certain computational problems cannot surpass the corresponding theoretical lower bounds.
- The Free Group Factor Problem: The paper gives a result covering all interpolated free group factors with parameter greater than 1, including the infinite-parameter case, and is accompanied by Lean formalization materials.
The repository README also discloses two exceptions: the Riemann zeta function manuscript and the Hodge conjecture manuscript are the only two results that deviate from the fixed process, and the Riemann zeta function write-up was human-edited to improve readability.
Lean Formalization: A Verification Benchmark, but Limited Coverage
In this release, Lean formal proofs are the core basis for technical credibility. Lean is a programming language that allows mathematical proofs to be checked step by step by computer—it does not accept "obviously true" leaps, and every step of derivation must be verified through logical rules.
According to public directory counts, about 162 of the 722 manuscripts have main conclusions accompanied by Lean formalization materials; the rest have not completed computer verification. OpenAI explicitly acknowledges in the README that some unformalized results may have problems and promises ongoing revisions. The repository contains only 10 summaries of model reasoning processes, rather than the full reasoning chain, meaning researchers cannot systematically trace the generation logic of each conclusion.
Daniel Halpern-Leistner, an associate professor of mathematics at Cornell University, is developing automated formalization tools. He previously told the media: "This now feels very urgent, because arguments written in natural language are pouring in." He also pointed out a key boundary: the machine checks the proposition defined in the code; if there is a deviation between the code definition itself and the original mathematical problem, passing formalization cannot endorse the entire paper. The Lean note for the Quasi-Riemann result explicitly states that the later application sections of the paper are not covered by that formalization.
This is not a rejection of Lean formalization, but a precise description of its boundaries: it is the strictest verification mechanism for AI mathematical output to date, but there is still a substantial distance between "the main conclusion has Lean" and "all conclusions in the paper are trustworthy."
AGMAI: Academic Governance Forced to Intervene
Behind this release is a governance mechanism that has received relatively little attention—AGMAI (Advisory Group on Mathematics and Artificial Intelligence), hosted by the Institute for Advanced Study (IAS), with members including Fields Medalists Timothy Gowers, Martin Hairer, and Edward Witten, among other leading mathematicians.
AGMAI's creation was itself reactive. In early September of this year, OpenAI was criticized by the academic community over the way it released results related to Navier-Stokes, and only afterward did it partner with IAS to form this advisory mechanism. AGMAI officially issued its responsible release recommendations on September 29. The recommendations came from a systematic review of more than 600 responses from the mathematics community, and their core contents include: when AI organizations release mathematical results, they should disclose the model name, prompts, reasoning summaries, and computational cost, and archive the papers in "an academic repository not controlled by any AI company."
AGMAI also specifically warned: mathematical results should not be released as a marketing tool to promote models, saying this practice "has already caused significant harm to the mathematics community."
This warning did not come out of nowhere. According to media reports, in August this year, OpenAI convened about 40 mathematicians for a closed-door meeting, suggesting the model had solved hundreds of difficult problems. Some participants at the meeting believed the company had promised not to release them all at once; however, an OpenAI spokesperson later said they were not aware of such an assurance. The two sides do not even agree on whether such a commitment existed.
The October 6 release stands in sharp contrast to the August controversy, both in scale and in transparency standards. OpenAI set up a version control mechanism in the repository, recording all revision histories, keeping old versions accessible, and providing a BibTeX citation format for each manuscript—a standard feature of academic norms, but not common in research releases led by AI companies.
After Standard Benchmark Saturation
OpenAI explained the evaluation context for this batch of manuscripts in the README: model performance on existing mathematics benchmark sets has approached saturation, forcing the evaluation team to include more open mathematics research problems in the test set. This is a structurally significant signal in the field of AI capability tracking—when standardized benchmarks can no longer distinguish top models, how to design credible frontier evaluations has itself become a methodological problem that needs to be solved.
The 722 manuscripts released this time essentially provide a new kind of reference frame for this problem: a set of real open mathematics problems that is reproducible, citable, and partially formally verifiable. The 372 problem families are classified by mathematical branch, and each family contains a main result, supporting arguments, corollaries, and alternative proofs, allowing researchers to verify from different angles. This is qualitatively different from previous ways of demonstrating AI mathematical ability, which were mainly through competition problem sets or closed evaluations.
Northwestern University mathematician Bryna Kra said in an interview with WIRED before the release: "This is an unsettling moment, but also genuinely exciting."
Independent Judgment
The significance of this release lies not only in the value of the mathematical results themselves, but in the precedent it sets: AI organizations can, under an academic advisory framework, publicly release the reasoning output of frontier models in a citable, version-traceable, and formally verifiable way. If this mechanism can be replicated by other AI organizations, it will establish a structural pathway for the academic credibility of AI output.
There are two points to watch in this release. First, the coverage gap in Lean formalization is substantial: about 560 of the 722 manuscripts have not yet completed computer verification, and it cannot currently be ruled out that the unformalized results have systemic problems. Second, AGMAI's call to archive in an "independent academic repository" has not been fully implemented—OpenAI chose a GitHub repository it controls. Technically, this meets openness requirements, but in terms of institutional independence, it still falls short of AGMAI's ideal standard.
Reproducible, verifiable AI mathematical output now has its largest public benchmark as of today. This is real technological progress. The next question is how quickly the mathematics community can complete independent evaluation of the 372 problem families—that speed will determine whether these manuscripts are ultimately accepted by the academic community as valid knowledge.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接