Terence Tao Criticizes OpenAI for Releasing 722 Mathematical Manuscripts at Once: Solving Problems Is Just a Tool, Mathematics Is Losing Fertile Ground

OpenAI has published 722 model-generated mathematical manuscripts on GitHub, prompting Fields Medalist Terence Tao to warn that AI's focus on solving probl

On October 6, 2026, OpenAI pushed 722 mathematical manuscripts generated by an internal model in bulk to a public GitHub repository. The manuscripts span pure mathematics and theoretical computer science, organized into 372 result families; the model's training began on August 28, 2026. Fields Medalist Terence Tao spoke publicly: AI aims to optimize problem solving, but the methodology, failed paths, derivative questions, and other byproducts bred during problem solving all disappear. He said "Math 1.0" has been optimized to an unsustainable critical point, and "Math 2.0" must redefine mathematical progress.

Tao had previously participated in the Mathematics and AI advisory group established by OpenAI on September 21. In his advisory capacity, he described the mission: "We currently face a very specific challenge—advising OpenAI on how to coordinate the release of a large number of important mathematical results, which are reportedly generated by its internal model." The existence of the advisory group shows that the mathematical community and OpenAI have not yet reached consensus on the timing and manner of releasing results.

AI Turns Problem Solving into an Assembly Line, but Bypasses the Process That Actually Matters

Tao's criticism must be understood within the internal logic of mathematical knowledge production. In the mathematical community, the value of an open problem lies not only in its answer. It is first a landmark: solving it tests human understanding of some piece of mathematical terrain. After a problem is solved, the mathematical community usually continues to work around the result for months or even years: understanding the structure of the proof, identifying new methods, connecting it to existing systems, simplifying it to the point where it can be taught, and eventually, decades later, some ideas become ordinary tools. This digestion process is a core output of mathematics.

AI's bulk problem solving interrupts this cycle. The 722 manuscripts were not released through traditional peer review; they were directly generated by the model and then pushed out. Of these, only 162 can be checked in the formal language Lean; the remaining 560 still require traditional rigorous scrutiny by the mathematical community. Tao's judgment: solving problems is merely a tool for achieving mathematical goals; the real goal is conceptual understanding and insight. If AI produces only "true/false" assertions without offering paths of thought, mathematics loses the fertile ground that cultivates new problems and methods.

He extends this into a call for the "Math 2.0" era: mathematical progress needs to be measured more comprehensively, and quality of exposition, community building, and the opening of new directions should be as important as solving problems. Otherwise, AI merely replaces an old unsustainable optimization with a new form of bulk compute consumption.

Navier-Stokes Controversy: A Flashpoint That Brought the Issue to a Head

The direct trigger for the discussion was OpenAI's September 8 announcement that it had solved the existence and smoothness problem for the Navier-Stokes equations. This is one of the seven Millennium Prize Problems established by the Clay Mathematics Institute in 2000. OpenAI deployed about 10,000 AI agents working continuously for about 88 hours, giving a proof that constructs a smooth external force causing the fluid to blow up in finite time, corresponding to cases C and D in the Clay problem statement. As of now, the official page of the Clay Mathematics Institute still marks the problem as "active."

The gap between the result's level of scrutiny and the standards of the mathematical community exposes a structural fault line between AI mathematical results and accepted standards. In mathematics, checking is a social process: peers read, question, simplify, and reproduce, eventually forming consensus. OpenAI used "automated methods for checking," which the company says are becoming a new standard for mathematical rigor, but this and the kind of checking the mathematical community expects are two different systems.

The Navier-Stokes incident also sparked disputes over attribution: New York University mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge were drawn in, and questions over whether the model brute-forced the problem while knowing about mathematicians' research progress, and whether the attribution of the result was reasonable, triggered public discussion in the mathematical community over whether a prompt counts as unpublished research material. This touches a deeper issue: top models receive users' conversation data, so how much of a proof produced by a model borrows from mathematicians' half-finished ideas?

25 Fields Medalists: A Systemic Misalignment Between AI and the Goals of the Scientific Community

On September 11, 2026, Terence Tao, Yu Deng, and 23 other Fields Medalists issued a joint statement, "The Serious Misalignment of Artificial Intelligence in Mathematics." The statement characterizes the problem as systemic: AI companies use "solving mathematical problems" as a benchmark for measuring model capability, which harms the discipline of mathematics and the mathematical community and has already affected other scientific and creative professions.

The statement lists three specific structural problems. First, AI solves problems faster than the mathematical community can digest the knowledge; the stock of open problems is being rapidly consumed, but AI does not yet have the ability to pose new open problems of far-reaching value. Landmarks are disappearing, and new landmarks are not being produced fast enough to keep up with consumption. Second, AI's way of releasing results has compliance problems: results are often hastily announced to the media, without rigorously written papers or proper citation of prior work, giving rise to attribution and plagiarism disputes. Third, if mathematicians are not responsible for integrating AI-generated ideas into the knowledge system, those ideas cannot truly "come alive," and the chain of transmission among human mathematicians may break.

Layered Impact on the Mathematical Community, the AI Industry, and the Evaluation Ecosystem

For the mathematical community, the impact of the 722 manuscripts is multifaceted. First is resource crowding: mathematicians must spend a great deal of time scrutinizing AI-generated claims, time that could otherwise be used to advance research. Second is a chilling effect: the incident has caused researchers to fear that large companies monitor prompt data in the background, steal inspiration, and race ahead with compute. Researchers are beginning to hesitate about revealing half-formed ideas in a chat box. If the mathematical community moves from open collaboration toward defensive, closed-door work, it would be a serious regression for the discipline's ecosystem.

For the AI industry, the incident brings a latent danger to the surface: as mathematical benchmark scores keep rising, what exactly is that score measuring? When a model deploys 10,000 agents to exhaustively check, score improvements come mainly from compute investment rather than essential advances in reasoning ability, and the effectiveness of mathematical problems as a capability evaluation tool is greatly reduced. Moreover, OpenAI's batch release "did not include all the disclosures recommended by the independent mathematics advisory group," meaning that even the advisory mechanism OpenAI itself established did not have its recommendations fully adopted.

For enterprise users and developers, this controversy provides a reference framework for selection: when deploying AI to solve professional-domain problems, there is an essential gap between "the model claims to have solved the problem" and "the solution has been validated by that domain's community." The distinction between the 162 manuscripts checked in formal Lean and the remaining 560 is a quantitative manifestation of that gap. In high-risk decision-making scenarios, relying on AI output that has not been checked by domain experts means assuming systemic risks that are difficult to quantify.

Strategic Judgment: The Credibility Crisis of Mathematical Benchmarks Will Force a Reconstruction of Evaluation Systems

Tao's "Math 2.0" framework is essentially a philosophical challenge to the AI capability evaluation system: when the proxy metric of "problem-solving rate" is optimized close to its ceiling, it is no longer a reliable signal of true capability and may even begin to diverge from true capability. This logic applies not only to mathematics but to all fields benchmarked by "solving checkable problems."

One most likely development is that mathematics will be the first field to form independent standards for checking AI results, requiring publishers to provide formal proofs (such as in Lean or Coq format) rather than natural-language manuscripts, and requiring the checking work to be done by independent mathematicians rather than relying on the publisher's own "automated checking." Observable signals include whether the Clay Mathematics Institute ultimately recognizes the Navier-Stokes problem as solved and whether OpenAI adjusts how it releases its next batch of results.

Another direction is for Tao's "Math 2.0" to move from an idea to institutional design: no longer measuring AI's mathematical value by "how many problems it solved," but by "how much understanding it produced that can be digested, transmitted, and taught." This would require the mathematical community, AI companies, and academic institutions to adjust their incentive structures simultaneously—extremely difficult—but the very fact that 25 Fields Medalists issued a joint statement shows that this path is being seriously discussed.

An underlying contradiction is this: AI companies need to prove model progress by solving famous problems, an inevitability of commercial communication logic; but the mathematical community needs to slow down to truly benefit, an inevitability of knowledge production logic. The time scales of these two logics are completely different, and the current pace of the AI industry is clearly dominated by the former. Tao's warning indicates that the cost of this pace is accumulating and has already attracted the attention of the field's most elite group.