10,000 Agents Crack Navier-Stokes in 88 Hours, but the Wrong Version—Controversy Overshadows the Breakthrough

OpenAI announced an AI-assisted proof of finite-time blowup for a forced version of the 3D Navier-Stokes equations, but the result does not solve the Clay Millennium Prize problem. The release has triggered a deeper dispute over authorship pressure, research data handling, and the ethics of AI-assisted scientific competition.

On September 8, 2026, OpenAI announced that its internal model had mobilized about 10,000 AI agents to complete a proof of finite-time blowup for the three-dimensional Navier-Stokes equations within 88 hours. Released alongside the announcement were a 166-page mathematical manuscript and a corresponding Lean formal verification project. Research lead Mark Chen told media that the computation costs alone were “as high as several million dollars.”

The Navier-Stokes equations describe fluid motion and form the mathematical foundation for aerospace engineering, climate simulation, and turbulence research. The question of existence and smoothness of their solutions has remained unresolved since the equations were fully formulated in 1845, and the Clay Mathematics Institute lists it among the seven “Millennium Prize Problems,” with a prize of $1 million. As of publication, the Clay Institute’s official website still lists the Navier-Stokes problem as unsolved.

First, Get the Math Straight: What OpenAI Proved, and What It Did Not

OpenAI itself acknowledged a crucial detail: the proof concerns the forced Navier-Stokes equations, meaning an external forcing term is introduced into the equations. OpenAI’s official blog explicitly stated that it does not intend to apply for the Clay prize, precisely for this reason—the Millennium Prize requires solving the freely evolving version without external forcing.

This distinction is not mathematicians’ pedantry. Adding external forcing is equivalent to installing a continuous “pump” into the fluid, fundamentally changing the energy structure of the equations; the technical difficulty is not on the same level as the unforced case. OpenAI’s proof shows that, under artificially imposed external forcing, an initially smooth velocity field can reach infinite velocity in finite time—that is, a “blowup” phenomenon. This is a real mathematical advance, but describing it as “conquering Navier-Stokes” involves an obvious leap in precision.

The operational scale of the agent cluster itself is impressive: according to reports, over 88 hours, these roughly 10,000 agents sent a total of 2.7 million messages and consumed about 130 billion output tokens. The model used was described as an internal version with capabilities “significantly exceeding GPT-6 Astra.” OpenAI developer Noam Brown cited a comparison: in 2025, Olympiad-level mathematics required OpenAI and Google DeepMind to expend enormous compute, while by 2026, a user with a $20 ChatGPT subscription could match the same level.

Another Team: A Year Earlier, Taking a Different Route

One day before OpenAI’s announcement, on September 7, New York University mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge released a set of preprints proving finite-time blowup for the three-dimensional Euler equations, including a forced version, accompanied by a 112-page paper and a public Lean formal verification. The Euler equations are closely related to the Navier-Stokes equations, with the latter degenerating into the former after a viscosity term is added. This work took about a year to complete and used AI tools including Claude, OpenAI Codex, and Astra throughout the process.

The mathematical routes chosen by the two teams were highly similar—Buckmaster himself said that “almost no one else was working” along this path. OpenAI also acknowledged in its announcement that its solution was accelerated after hearing “rumors” about Anthropic’s direction. According to reports, OpenAI mathematics team lead Sébastien Bubeck said the company began training a new advanced mathematics model on August 28, allocated more resources to the Navier-Stokes problem after hearing the relevant rumors, and ultimately “obtained the complete Lean-formalized solution on Sunday morning.”

The Real Storm: Authorship Pressure and Data Provenance

The core of the controversy is not the mathematics itself, but what happened during the research process.

Buckmaster publicly disclosed that on September 6, OpenAI contacted him, claiming that its internal model had produced a roughly 100-page proof using an approach “strikingly similar” to the work of Buckmaster and Alpöge. Bubeck then twice proposed publishing jointly—but explicitly demanded that Alpöge be removed from the author list on the grounds that he was an Anthropic employee. Bubeck described the situation as “so annoying.” When Buckmaster refused, Bubeck allegedly said, “Why are you trying to ruin your career?” and “If you don’t want me to be nice, I can also not be nice.”

In response to those claims, Bubeck told WIRED that what he had discussed was inviting Buckmaster to lead the rewriting of OpenAI’s Navier-Stokes proof, and that he considered Alpöge’s involvement in the project inappropriate because Alpöge was an employee of competing company Anthropic. He admitted that he did use the wording about “career” and apologized for it, saying he immediately withdrew the remark. OpenAI’s official statement said the company’s researchers “had not seen their work through any channel until they publicly released it last night.”

However, this defense faces a technical question that Buckmaster has raised publicly: throughout the research process, he and Alpöge uploaded all drafts into conversations with OpenAI Codex. Were those session data used to train later models? OpenAI’s answer under questioning was “unclear,” which has not fully convinced outside observers. OpenAI has not directly explained how those session data were handled.

A Structural Contradiction: Both Infrastructure and Competitor

This controversy is not merely a dispute over which team came first. It reveals a fundamental contradiction in AI-assisted research today: AI labs are simultaneously playing two conflicting roles—providers of research infrastructure and institutions actively competing in scientific research.

Buckmaster and Alpöge used Codex as a research tool, which is a reasonable and normal choice today, much like physicists using computing clusters or statisticians using the R language. But when the tool provider is also a potential beneficiary of intellectual property, the boundaries around research data become blurred: are drafts “user data” or “training material”? Can conversation records enter model training under compliant terms? Under OpenAI’s existing user-agreement framework, these questions are not free from dispute, and academia has long lacked clear norms for such scenarios.

Bubeck’s demand to remove an Anthropic employee from co-authorship—whatever the motivation—brought this contradiction into the open. When authorship of a mathematics paper becomes tied to the competitive interests of an AI company, the basic principle on which the academic community operates—that credit for results should be determined by scholarly contribution—comes under direct attack. Should an Anthropic employee be barred from authorship on a mathematical proof completed using OpenAI tools simply because of his employer? There is currently no industry consensus on this question, but it clearly should not be decided unilaterally by the head of a lab during a pressure call.

The Logic and Cost of a Compute Race

From a purely commercial standpoint, OpenAI’s actions have an internal consistency: the investment of millions of dollars in compute and the scheduling of a 10,000-agent cluster send a clear signal—that in the new battlefield of AI-assisted research, being the first to command attention means having pricing power. Noam Brown’s remark is worth parsing carefully: he framed the trend of reducing problem-solving costs from enormous compute to a $20 subscription as a positive narrative of democratized capability. That is certainly real, but in this specific incident, the same narrative also obscures an ongoing fact: the advantage of scaled compute allows a single institution to “follow up” on another party’s research direction in an extremely short time and announce first.

Buckmaster and Alpöge spent about a year; OpenAI used less than two weeks after hearing the rumors. This time gap itself is a direct manifestation of compute inequality. When the pace of mathematical discovery is compressed from “human scholars’ years” to “agent clusters’ days,” the determination of who came first will increasingly depend on the timing of institutional announcements rather than the actual sequence in which the work was completed.

Independent Assessment

The mathematical result OpenAI released this time is a real mathematical advance, but the narrative frame of “solving the Navier-Stokes Millennium Problem” is misleading—OpenAI knows this itself, which is why it proactively stated that it would not apply for the Clay prize. The more serious problem lies not in the mathematics itself, but in the competitive process.

If Buckmaster’s account is true, then this incident means that the research lead of an AI company, after hearing rumors of progress by external competitors, urgently followed up and then attempted to trade a career threat for concessions on community norms before public release. Whatever the ultimate legal and ethical judgment, this pattern of behavior itself will cause long-term damage to the credibility of AI tools in academic research. Will researchers still upload working drafts to commercial AI platforms in the future?

OpenAI’s current response—“we had not seen their work”—may be technically true, but it avoids the more central issue: why has there still been no clear answer about whether Codex conversation data were used for training? The silence on data provenance is more worthy of vigilance than any statement.

What the independent mathematics community needs to do is conduct rigorous peer review of the two teams’ 166-page and 112-page proofs respectively. Until then, there is no answer to the question of “who solved Navier-Stokes first”—but the question of “where the ethical boundaries lie in AI research competition” has already become impossible to avoid because of this incident.