OpenAI Claims AI System Produced a Lean-Formalized Navier–Stokes Disproof

OpenAI says an internal AI system found a finite-time singularity in Navier–Stokes equations, but independent mathematician review remains essential.

AI News

OpenAI says an internal AI system has produced an analytical proof that three-dimensional incompressible fluid motion can develop a finite-time singularity, addressing the Navier–Stokes existence and smoothness problem. The company has released a written argument and a formalization in Lean, but the claim has not been independently validated in the available evidence.

The announcement matters because the Navier–Stokes problem is one of the seven Millennium Prize Problems, a group of mathematical challenges identified by the Clay Mathematics Institute in 2000. A correct solution would have implications for the limits of automated mathematical research as well as for how AI systems are tested on problems that have resisted human researchers for decades.

OpenAI’s post is the principal source for the account. The accompanying wire item contains no additional article text in the available material, so the strongest claims about the system’s performance, scale, and result are vendor-reported.

What OpenAI says it solved

OpenAI says its system established the version of the problem in which a smooth, initially resting fluid develops a singularity in finite time, while remaining subject to a smooth external force and retaining finite energy. In the company’s description, the result proves statements “C” and “D” in the official formulation of the Navier–Stokes problem, which would amount to a disproof of universal smoothness rather than a proof that all solutions remain regular.

The proposed mechanism is a vortex that spirals inward and stretches along its axis. As the central region contracts, the fluid’s speed increases without bound in the mathematical model. OpenAI says the construction depends on the acceleration, pressure, momentum-transfer, and viscosity terms becoming large while canceling in a precise way. The external force remains smooth, meaning the claimed breakdown is generated by the fluid dynamics rather than by inserting an infinite force directly.

The Navier–Stokes equations model fluids as continuous media using Newton’s laws of motion. Engineers and scientists use them in areas including aircraft design, weather forecasting, and blood-flow analysis. The unresolved issue is whether a smooth three-dimensional incompressible flow with constant density must remain smooth for all time.

A large multi-agent experiment

According to OpenAI, the work began on September 1 after the company heard rumors that two Millennium Prize Problems had been resolved. It then directed a coordinating system of AI agents to examine all of the open problems and several related questions.

The company says the Navier–Stokes effort involved roughly 10,000 concurrent agents. Those agents were divided into groups, given different formulations of the problem, and allowed to use tools including a cached version of the internet and code execution. OpenAI says the agents explored both proof and disproof variants before the most promising work was consolidated.

OpenAI also reports that the broader project generated 4.9 million messages and used about 300 billion output tokens. For the Navier–Stokes effort specifically, it reports 2.7 million messages and approximately 130 billion output tokens. These figures describe the company’s internal experiment and are not an independently audited measure of useful mathematical reasoning.

The company says agents reached the Navier–Stokes result on September 5, after about 88 hours of work. Lean formalization and verification took another 17 hours through GPT-6 Astra, according to OpenAI. The company says Codex was used to consolidate insights from different agent groups during the search.

Formal verification is evidence, not final acceptance

The Lean artifact is an important part of OpenAI’s disclosure because formalization can expose missing assumptions and invalid steps that may survive in a conventional research draft. It does not, by itself, settle whether the formalized statement correctly matches the official Millennium Prize formulation or whether the implementation contains an error in definitions, imported lemmas, or hypotheses.

OpenAI has published the proof writeup and Lean formalization, but the supplied evidence does not include a review by the Clay Mathematics Institute or independent mathematicians. The company explicitly says it does not intend to claim the Millennium Prize for the result. That distinction is significant: a public AI-generated proof is a research claim, while prize recognition would require scrutiny by specialists and acceptance under the institute’s rules.

The timing also creates a need for careful attribution. OpenAI says it began the work after hearing rumors connected to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a mathematics professor at New York University. The company says their work concerned forced Euler rather than the same Navier–Stokes result, and that the two efforts differed in both methods and precise conclusions. OpenAI says it did not access their specific user data or work before public release, while acknowledging that it cannot rule out de-identified product data having helped improve its models.

What the result could mean for AI builders

For AI developers, the notable change is not simply that a model generated mathematical text. OpenAI describes a workflow in which many agents independently explore formulations, exchange intermediate results, use software tools, and hand a candidate proof to a formal verification system. That architecture resembles an automated research pipeline more than a single chatbot response.

The approach also highlights its costs and operational risks. Hundreds of billions of generated tokens and thousands of concurrent agents may be practical for a high-value research benchmark, but they are far removed from the economics of ordinary software products. Builders would need to measure whether parallel exploration produces genuinely new insights or mainly increases the volume of plausible but redundant reasoning.

For enterprise buyers, the episode reinforces the difference between a model’s ability to propose a result and an organization’s ability to trust it. In scientific and engineering workflows, reproducible artifacts, machine-checkable proofs, expert review, and clear records of assumptions matter more than a benchmark score or the scale of an agent run. The Lean formalization may support that process, but it does not remove the need for domain expertise.

The announcement is also a competitive signal. OpenAI says the internal model is more capable than GPT-6 Astra and that its training was still continuing. Because those capability comparisons come from OpenAI, they should be treated as internal or vendor-reported claims rather than settled market measurements.

What to watch next

The first signal will be independent examination of the published proof and Lean code, especially whether the construction satisfies every condition in the official Navier–Stokes statement. A formal proof that depends on a subtly different domain, force condition, or notion of singularity would not resolve the same problem.

The second is whether mathematicians reproduce the result, identify gaps, or confirm that the formalization compiles and faithfully encodes the argument. The response from the Clay Mathematics Institute would also clarify whether the work meets the standards for recognition.

Researchers should also watch whether OpenAI releases enough implementation detail to reproduce the multi-agent search, including model versions, prompts, tool boundaries, and compute accounting. Without that information, the reported agent counts and token totals provide context but limited guidance for teams building similar systems.

Creati.ai perspective

OpenAI’s announcement is best understood as a significant capability claim, not yet a settled mathematical fact. The combination of agent coordination, code tools, and Lean formalization is a credible direction for AI-assisted research, but the central result still depends on scrutiny outside the company that produced it.

For builders, the practical lesson is narrower and more useful than the headline: advanced AI research systems need verification layers, explicit problem specifications, and independent review. If those controls hold up around this result, the episode could become a meaningful case study in machine-assisted mathematics. If they do not, it will show why fluent proofs and formal-looking artifacts cannot substitute for mathematical validation.

Ads