📑 Table of contents

OpenAI solves Navier-Stokes in 88 hours with 10,000 agents: public Lean proof, prize declined, authorship controversy

Actu IA 🟢 Beginner ⏱️ 11 min read 📅 2026-09-09

OpenAI solves Navier-Stokes in 88 hours with 10,000 agents: public Lean proof, prize declined, authorship controversy

🔎 A 180-year-old problem falls in a single weekend

On September 5, 2026, OpenAI announced it had solved one of the seven Millennium Prize Problems: the behavior of solutions to the Navier-Stokes equations. An unreleased internal model, orchestrating up to 10,000 AI agents working in parallel, produced in 88 hours a proof of finite-time blow-up — the core of the mathematical challenge posed by the Clay Mathematics Institute in 2000, with a $1 million prize attached.

The demonstration was not kept under wraps: complete write-up, public PDF, and a verifiable Lean formalization on GitHub, confirmed in 17 additional hours by GPT-6 Astra. And yet, no one will be collecting the check. OpenAI refuses to claim the prize, and the mathematical community, for its part, is demanding answers.

Because behind the feat lies an explosive authorship dispute. Tristan Buckmaster (NYU) and Levent Alpöge claim to have been working for months on exactly the same approach — and suspect that their own work sessions were used to train OpenAI's models. The question now goes beyond the media splash: what is the value of a machine-verifiable proof, when its provenance cannot be verified?

The Essentials

  • 88 hours, 10,000 agents: an unpublished internal model from OpenAI solved the finite-time blow-up of the Navier-Stokes equations, one of the 7 Millennium Prize Problems (worth $1M each since 2000, according to MIT Technology Review).
  • Publicly verifiable proof: write-up, PDF, and Lean formalization on GitHub, verified in 17 hours by GPT-6 Astra. OpenAI refuses to claim the $1M prize from the Clay Institute.
  • Colossal cost: roughly 300 billion output tokens, i.e., ~$22.5M of compute at current rates — about 1000 times more than previous AI mathematical results (Sébastien Bubeck, cited by TechCrunch).
  • Authorship controversy: Buckmaster (NYU) and Alpöge suspect a leak from their research sessions; OpenAI denies it but "cannot rule out" that de-identified data improved its models.
  • The deeper question: mechanical verifiability of a proof says nothing about its ethical origin. Science is entering the era of uncertain provenance.

Tool Main Use Price (September 2026) Ideal For
GPT-5.5 (OpenAI) Mathematical reasoning, multi-agent orchestration Starting at $20/month (ChatGPT Plus, check openai.com) Researchers wanting to test formal approaches
Claude Opus 4.7 Proof analysis, mathematical writing $20/month (Claude Pro, check anthropic.com) Critical review of scientific papers
Gemini 3 Pro Deep Think In-depth documentary research Included in Google AI Pro (check gemini.google.com) Academic monitoring and source cross-referencing
Kimi K2.6 Self-hosted mathematical research Free as open weights (infrastructure costs) Labs concerned with data privacy

To build your own local research agent pipelines, check out our guide on open source AI agents with Ollama and our comparison of the best LLMs for AI agents.


What actually happened — the verified facts

Direct answer: an unreleased internal OpenAI model produced a proof of finite-time blow-up of the Navier-Stokes equations in 88 hours, with up to 10,000 agents working quasi-concurrently, and OpenAI has published the formal proof for public verification.

The figures are corroborated by several media outlets. According to CNN, the company itself stated that its model "took 88 hours to solve the problem, using up to 10,000 AI agents working somewhat concurrently" (CNN). The effort reportedly involved 2.7 million messages exchanged between agents and roughly 300 billion output tokens, i.e., $22.5M of compute at current rates (TechCrunch).

The problem, in two sentences

The Navier-Stokes equations describe the motion of fluids. The Millennium question asks to prove (or disprove) the existence and regularity of smooth solutions under all circumstances — or, as here, to demonstrate finite-time blow-up for the forced equations, that is, solutions that become infinite in bounded time.

What OpenAI has published

  • A detailed write-up of the proof.
  • A complete formalization in Lean, the proof assistant, verifiable by anyone on GitHub.
  • A cross-check: GPT-6 Astra confirmed the formalization in 17 additional hours.

It is this last point that makes the affair historic from a methodological standpoint: no matter who wrote the proof, the machine can verify it. This is the first time a Millennium Problem has fallen with a formal certification from the moment of the announcement.


Why OpenAI refuses the $1 million prize

Direct answer: because claiming the prize would require clear authorship, a Clay Institute peer-review process, and transparency about the method — three things OpenAI cannot guarantee without exposing itself.

The Clay Mathematics Institute traditionally requires publication in a peer-reviewed journal and acceptance by the community over a period of at least two years. Here, three obstacles stand in the way:

  1. Authorship is contested (see below).
  2. The method is industrial, not academic: 10,000 agents and $22.5M of compute (New Scientist) fit no category provided for in the rules.
  3. The model is unpublished: the community cannot examine how the proof was constructed, only what it is.

This refusal is as clever as it is embarrassing. OpenAI captures the prestige without facing the academic validation process — but it also fuels suspicion: why refuse the validation process if the proof is solid?


The controversy: Buckmaster, Alpöge and the shadow of a leak

Direct answer: two mathematicians had been working on the same approach for months, and one of them claims to have been informed of OpenAI's proof before even publishing his own.

On September 5, Tristan Buckmaster (NYU) and Levent Alpöge made public three results related to Navier-Stokes. In his official statement, Buckmaster writes that he learned "that an internal OpenAI model had produced a proof of finite-time blow-up for the forced Navier-Stokes equations" — even as he and Alpöge were using "a specific method on a related problem." The timing is damning: the scientific press reports that OpenAI acknowledges having been inspired by rumors that Alpöge and Buckmaster had solved a Millennium Problem.

The accusations, point by point

  • Suspicion of data leakage: Buckmaster and Alpöge suspect that their working sessions with AI tools were used to train or guide OpenAI's models. The Hindu mentions accusations of "unethical practices," including the alleged use of stolen data and an aspect related to privacy involving Sébastien Bubeck and Codex.
  • OpenAI's response: the company denies any intentional leak, but admits that it "cannot rule out" that de-identified data from its customers' usage may have improved its models — an admission which, legally speaking, says it all.
  • "OpenAI played dirty": that's the headline TechCrunch attributes to the NYU mathematician. Fortune reports accusations of cheating and intimidation, along with Terence Tao's "lament," deploring the way the announcement overshadowed human work.

My opinion, stated plainly: regardless of the outcome of the investigation. The mere fact that a private lab could, within a single week, swallow the equivalent of decades of postdocs' worth of computation — and cast a shadow over two careers built on this problem — should worry the entire scientific community.


$22.5M of Compute: The Price of Discovery

Direct answer: it's "emphatically millions of dollars," according to Sébastien Bubeck, or roughly 1000 times the cost of AI-achieved mathematical results to date.

Metric Value Source
Resolution time 88 hours CNN
Parallel AI agents up to 10,000 CNN / Anadolu Agency
Inter-agent messages 2.7 million Cross-checked from viral post
Output tokens 300 billion TechCrunch / cross-checks
Estimated compute cost ~$22.5M (current rates) TechCrunch
Lean verification 17 h (GPT-6 Astra) Cross-checks

For comparison, OpenAI's announcement on the Erdős problem — a geometry theorem that had resisted proof for 80 years — involved a radically lower cost regime. We've moved from feasibility demonstration to industrial-scale brute force: the question is no longer "can it solve X?" but "who can afford to solve X?".

And that may be the real message. Millennium Problems are now within reach of players able to spend $22.5M of compute in a single week. University labs, meanwhile, are watching from the sidelines.


Verifiable proof, unverifiable provenance: the new scientific dilemma

Direct answer: for the first time, the formal validity of a major result is beyond doubt, but its ethics of production cannot be established. Science must invent norms of provenance, not just of verification.

The Lean formalism changes the game irreversibly. A formalized proof is verified by a small logical kernel; there is no "almost correct," no referee's arbitration. As several observers summed it up: mathematically, the result either holds or it doesn't — and here, it holds.

But three tensions emerge:

  1. Provenance becomes the blind spot. A proof can be correct and contaminated: built on data obtained without consent, or on the ideas of an uncredited third party. The peer-review system has no tools for this.
  2. Competition replaces collaboration. The historical model — shared conjectures, seminars, preprints — collides with 88-hour races conducted in secret. Terence Tao is not the only one to lament this.
  3. Structural asymmetry. OpenAI "cannot rule out" that its users' data helped train its models. As long as that sentence remains possible, every mathematician using a commercial tool must wonder whether they are training their competitor.

What the community should demand

  • Strong contractual commitments that research sessions will not be used for training.
  • Provenance logs: which data, which interactions fueled the discovery.
  • A credit framework: if an AI proof builds on a researcher's published or private approach, co-authorship must be discussed, not swept aside.

Common Mistakes (in Media Coverage)

Mistake 1: "AI solved Navier-Stokes"

Not exactly. What was proven is a finite-time blow-up for forced Navier-Stokes equations — one of the scenarios eligible for the Millennium Prize, not the entire fluids research program. Applied fluid mechanics itself hasn't changed one bit this week.

Mistake 2: "The model proved the theorem all by itself"

The 10,000 agents are not a single brain: they are coordinated instances, exchanging 2.7 million messages, on an unpublished orchestration architecture. The exact role of the scaffolding versus the underlying model remains unknown — and that is precisely what the mathematicians calling for more methodological transparency are contesting.

Mistake 3: "OpenAI cheated, the proof doesn't count"

The suspicion about provenance doesn't change formal validity: the Lean proof is verifiable by anyone, and GPT-6 Astra independently confirmed it in 17 hours. Confusing production ethics with logical validity means missing both debates — and that's exactly what the most divisive coverage does.


❓ Frequently Asked Questions

Will the Clay Institute's $1M prize be awarded?

No, not in this form. OpenAI is not claiming it, and the Clay rules require publication, peer review, and a two-year community acceptance period. With an unpublished model and contested authorship, the conditions are not met — and likely never will be for this announcement.

Is the proof actually reliable?

Logically, yes: the Lean formalization can be verified by anyone and was confirmed in 17 hours by GPT-6 Astra. However, there is no guarantee that the proof is readable, minimal, or comprehensible to humans — a point mathematicians are already starting to examine.

Who are Buckmaster and Alpöge?

Tristan Buckmaster is a mathematician at NYU and a leading expert in fluid equations. Levent Alpöge is his collaborator. They made three related results public the same weekend and claim to have been working for months on the approach that OpenAI's proof implements.

Did OpenAI steal their data?

Nothing is proven. OpenAI denies any intentional leak but acknowledges it "cannot rule out" that de-identified data from customer usage improved its models. An investigation or independent audit would be needed to settle the matter — none has been announced so far.

What does this change for mathematics?

Three things: the Millennium Problems no longer protect researchers from industrial-scale brute force; formal verification (Lean) becomes a publishing standard; and the provenance of AI-assisted discoveries becomes a central ethical issue, on par with validity.


✅ Conclusion

In 88 hours and with $22.5M of computation, OpenAI transformed a two-century-old problem into a verifiable proof — and into a textbook case on the provenance of AI discoveries. The lesson to remember: in the era of machine-certified proofs, the question is no longer just "is it true?", but "where does it come from?". If you'd like to explore the multi-agent architectures that make these feats possible yourself, start with our guide to the best autonomous AI agents.