📑 Table of contents

OpenAI publishes proofs of 722 math problems solved by a secret model — academic research overtaken by AI

Actu IA 🟢 Beginner ⏱️ 14 min read 📅 2026-10-07

OpenAI publishes proofs of 722 mathematical problems solved by a secret model — academic research outpaced by AI

🔎 722 manuscripts out of nowhere, a model no one has ever seen

On October 6, 2026, at 21:47 UTC, OpenAI pushed to GitHub a repository modestly named openai/math. Inside: 722 mathematical manuscripts, organized into 372 families by discipline, produced by an internal model that no one outside the company has ever seen in action. No name, no prize, no API endpoint. PDFs, source files, citation instructions — and 162 proofs formalized in Lean, machine-verifiable.

Two months after the announcement of the Navier-Stokes solution, the company is scaling up: this is no longer an isolated feat, it's a catalog. Sam Altman speaks of a "new era of discovery" (Fox News, October 7, 2026), and the tech press — The Verge leading the way — has echoed a publication without precedent in the history of AI.

Why now? Because the sequence is accelerating. September 8: Navier-Stokes. October 6: 722 manuscripts. Each publication widens the gap between what AI produces and what the mathematical community can absorb. And this time, OpenAI isn't publishing a result — it's publishing a pipeline of results.

The problem boils down to one sentence: the community validates at human speed, the model produces at industrial speed. No peer review to date, and a warning from OpenAI itself: "some of the non-formalized results could be problematic."


Key takeaways

  • 722 manuscripts, 372 families: published on October 6, 2026 on GitHub under the Apache-2.0 license, with PDFs, source files, and citation protocols.
  • 162 proofs formalized in Lean, i.e., roughly 1 in 5 manuscripts with a machine-verifiable proof. For the other ~560: caution is required.
  • ~4,000 problems explored, each result costing on average ~3 hours of ChatGPT Pro-style "thinking" compute, using an internal model that was never released.
  • Zero peer validation to date. The Advisory Group on Mathematics and AI at the Institute for Advanced Study (Princeton): the release "is the beginning of a process, not the end."
  • OpenAI promises workshops and future access to the model, with no date given, and will not claim the Millennium Prize for Navier-Stokes.

Tool Main use Price (October 2026) Best for
openai/math (GitHub) Read the 722 manuscripts, sources, revision history Free (Apache-2.0) Mathematicians, reviewers
Lean Verify the 162 formalizations Free, open source Independent re-checks
ChatGPT Pro "Thinking"-class reasoning, manuscript review $200/month (check openai.com) Researchers, analysts
NotebookLM Digest of long PDFs, cross-referencing sources Free Monitoring and synthesis

To go further in choosing a research assistant — literature monitoring, manuscript synthesis, formalization — our guide Best AI for research compares what actually serves a researcher, beyond publisher marketing.


What OpenAI published — and what it carefully kept

OpenAI published manuscripts, not community-validated results. That distinction shapes the entire debate that follows.

The openai/math repository is under the Apache-2.0 license and its organization is that of a genuine scientific publication: a preprints/ folder containing the PDFs, the source files, and the citation instructions, plus a Lean library of 162 entries. About one manuscript in five therefore comes with a proof whose correctness can be checked by a machine — not by a tired reviewer.

The Apache-2.0 license is not a detail: anyone can fork, reuse, and re-verify. It's the minimum required for an independent re-check to be possible, and OpenAI has done so.

A detail worth noting: the revision history is public, earlier versions included, as confirmed by OpenAI's community post of October 7. The text devoted to the zeta function was also revised by humans for readability, according to Orcarouter's analysis. You can trace the evolution of a result — real transparency, and rare in this industry.

What is not published says just as much. The model remains nowhere to be found: no name, no technical report, no access. GitHub issues are disabled and pull requests are reserved for collaborators. The openness is one-way: everyone can read, no one can discuss.

Results already making history

The official PDF "Ten Advances in Mathematics and Theoretical Computer Science" lists the centerpiece results of the catalog. Three are enough to gauge the ambition:

  • Sphere packing: an improvement of the Cohn–Elkies exponent, the first general improvement since 1978. Nearly fifty years of waiting.
  • Connes' rigidity conjecture: refuted. A major open problem, whose refutation will make waves far beyond the theory of operator algebras.
  • Two Erdős conjectures: refuted. We detail in our article on the proof of a geometry theorem that had resisted for 80 years what this kind of result changes in concrete terms.

Added to these are the construction of an explicit non-sofic group and multicolor Ramsey bounds of k^Θ(k) for R_k(3) — advances that Mezha summarizes as a mix of solutions to open problems, new proofs, and refutations of established conjectures.

Family 003 — the inequality ζ(s) ≠ 0 for Re(s) > 7/8, a direct step toward the Riemann hypothesis — is already the subject of a public independent re-check, launched from the repository. This is exactly the intended scenario: the machine proof as the starting point, the community as the final arbiter.


The model behind it all: a ghost, fully owned

No one outside OpenAI can use this model, and the company won't say when that will change.

What we know fits in a few lines, drawn from the Navier-Stokes announcement of September 8, 2026: an internal model "significantly outperforming GPT-6 Astra", trained starting August 28, 2026 (Unite.AI). It explored ~4,000 problems, and each result cost on average ~3 hours of ChatGPT Pro thinking-style compute. In total, a few thousand hours of compute — a derisory budget for a lab, unattainable for a lone researcher.

Put that figure in perspective: a mathematician devotes years to a single problem. This model produced 722 manuscripts in a matter of weeks. The marginal cost of producing research has just collapsed, and no one has finished measuring the consequences.

The best public models — Gemini 3.1 Pro (92), GPT-5.5 (91) according to our ranking — excel on standard benchmarks. Not one has published a single original research manuscript. The gap is not one of degree, it is one of kind.

This silence about the model is a business decision, not an oversight. On the product side, OpenAI has just launched the preview of GPT-5.6 Sol amid a full-blown price war; a model capable of refuting Erdős conjectures has, for now, no place in an API billed by the token. The move upmarket had been foreshadowed when OpenAI made Path to Astra official, its first model to cross the critical threshold: research first, the product after.

And GPT-6 Astra, merely a point of comparison here, illustrates another issue: after SpaceX's acquisition of Cursor, OpenAI cut off access to the model with maximum notice, brutally reminding everyone who controls the infrastructure. Imagine an entire scientific field depending on a model that no one can audit.


"Impressed and worried": the mathematics community switches into forced review mode

Mathematicians are not disputing the feat: they are disputing the pace and the rules of the game.

OpenAI did not act alone in the sharing: it says it consulted the Advisory Group on Mathematics and Artificial Intelligence of Princeton's Institute for Advanced Study, as reported by ETV Bharat. The group delivered its recommendations on September 29, 2026, after more than 600 responses from the mathematics community: share with the actors in the field, fund human understanding of the results.

But the group also sets the condition that stings: publication "is the beginning of a process, not the end," and "humans must still" validate. Translation: 722 manuscripts is a review backlog, not an established library.

OpenAI itself warns, in the deposit, that "some of the non-formalized results could pose problems." Of the 722 manuscripts, about 560 have no machine-verifiable proof. And no result has passed peer review — the standard that has carried authority in mathematics for decades.

The reaction from the field is commensurate: "impressed and worried," in substance. Impressed by the density of the results. Worried by a machine that produces faster than the profession can review — and by a timeline imposed without prior consultation on the substance.

The workshops announced by OpenAI and the future access to the model are the counterpart offered to the community. No dates yet, though. In the meantime, the public re-check of the 003 family is moving forward on the forum: the reviewing has begun, at full scale.


From Navier-Stokes to 722 manuscripts: 38 days to change scale

The October 6 publication is not an isolated event: it is the culmination of a sequence orchestrated since late August.

Date Event
August 28, 2026 Start of internal model training
September 1, 2026 Launch of the Millenium Prize evaluation campaign
September 5, 2026 Resolution of Navier-Stokes (finite-time singularity, formalized in Lean)
September 8, 2026 Public announcement by OpenAI
September 29, 2026 Advisory Group recommendations (600+ community responses)
October 6, 2026, 9:47 PM UTC Publication of the 722 manuscripts on GitHub
October 7, 2026 Community post and launch of the public re-check of family 003

Thirty-eight days between the start of training and the publication of a catalog of this size. For comparison, the proof of a major conjecture traditionally occupies a mathematician — or a team — for years, publication and peer review included.

After Navier-Stokes, OpenAI had already laid out the next phase of its program, including the Hodge conjecture, as we reported in After Navier-Stokes, OpenAI targets the Hodge conjecture and tries to defuse the controversy. The October 6 message is more radical: the model does not target famous problems one by one, it produces continuously.

One ethical point deserves note: OpenAI will not claim the one-million-dollar Millennium Prize (Clay Mathematics Institute) attached to Navier-Stokes. The rules require a publication accepted by the community and two years without contestation — a process that the blunt release of 722 unreviewed manuscripts does not satisfy. Forgoing the prize also means dispensing with that scrutiny. A shrewd trade-off, and a revealing one.


The real divide: producing has become cheap, validating has not

For the first time in mathematics, the bottleneck is no longer the generation of results, but the human capacity to verify them.

The figures in the publication say it bluntly. 722 manuscripts in 38 days — that's more than the annual output of several mathematics departments combined (our own estimate). 162 proofs formalized in Lean: several years of work for the most active formalization teams in the world. The imbalance is not a detail; it is the structure of the problem.

Peer review, meanwhile, runs at human speed: months for a paper, years for a major result. The re-check of family 003 — a single result out of 722 — is already drawing volunteers on a public forum. Do the math: at this rate, the catalogue won't be "validated" any time soon.

The consequences extend beyond mathematics. Who gets the credit when a model refutes a conjecture on which an entire thesis rested? How do you cite a manuscript revised by humans "for readability"? What editorial policy should journals adopt when facing 722-piece submissions? The Advisory Group is calling for funding of human understanding of the results: that is an official admission that verification has become the scarce resource.

And the economic context adds another layer. On October 7, Ray Dalio warned that the AI bubble could burst (Fox News). In this climate, a demonstration of frontier-level research capability is not just a gift to mathematicians: it's an investment argument. Worth keeping in mind when reading the announcements.


How to follow this work without getting trapped

Start with the 162 Lean formalizations, and treat everything else as conjectures until independently verified.

Four reflexes, in order:

  1. Prioritize what's verified. The repository's Lean library contains 162 entries whose correctness is machine-checkable. It is the only subset that can be cited without reservation today.
  2. Read the revision history. Earlier versions are accessible, and some have changed — the text on the zeta function was rewritten by humans. Cite the version, not the title.
  3. Follow the public re-checks. The OpenAI forum thread on the 003 family documents independent verification in real time. This is the model worth supporting.
  4. Beware of secondhand sources. Many relays turn "published" into "proven". Go back to the primary sources: the GitHub repository, official PDFs, the forum.

To gauge how public models measure up against this level of reasoning — a GPT-5.5, a Gemini 3.1 Pro, a Claude Opus 4.7, which dominate our rankings — our guide Best LLMs for Research is updated continuously. No publicly accessible model publishes original research manuscripts today; the gap with OpenAI's internal model remains total.


❌ Common Mistakes

Mistake 1: mistaking "published" for "proven"

722 published manuscripts, 162 machine-verifiable, zero peer review to date. OpenAI itself warns that unformalized results "could pose problems." The solution: only cite Lean formalizations, or results explicitly confirmed by a documented independent re-check.

Mistake 2: believing OpenAI has "solved the Millennium Problems"

Only one concerns the seven prize problems: Navier-Stokes, announced on September 8, 2026 — and OpenAI will not claim the million dollars. The 722 manuscripts cover other families of problems, from Erdős conjectures to Connes rigidity.

Mistake 3: looking for the model in the API or in ChatGPT

The model has "no name, no price, no API endpoint" (Orcarouter, October 2026). GPT-5.6 Sol is an entirely different, commercial product. Confusing the two leads to attributing to consumer subscriptions capabilities that remain, for now, locked away internally.

Mistake 4: citing a manuscript without checking its version

The revision history is public, and texts have been modified — sometimes by humans, for readability. Before any sharing or educational use, check the dated version of the PDF in the preprints/ folder of the repository.


❓ Frequently Asked Questions

Which model produced these 722 manuscripts?

An internal frontier model from OpenAI, never unveiled: no name, no price, no API endpoint. According to the Navier-Stokes announcement of September 8, 2026, it is "significantly more capable than GPT-6 Astra." Trained starting August 28, 2026, it explored roughly 4,000 problems, at a rate of ~3 hours of compute per result obtained.

Are the 722 proofs reliable?

162 out of 722 have a machine-verifiable Lean formalization — about one in five. The rest await human review, which has not yet taken place at scale. OpenAI itself warns that "some of the unformalized results could pose problems," and no result has passed peer review.

What is Lean, and why does it matter?

Lean is an open source proof assistant that verifies the correctness of a demonstration mechanically, without human intervention. A proof formalized in Lean eliminates review errors. Of the 722 manuscripts, 162 have one — it is the only subset that can be cited without reservation as of October 2026.

Will OpenAI receive the Millennium Prize for Navier-Stokes?

No, the company has said so itself. The Clay Mathematics Institute's rules require publication in a peer-reviewed journal and two years without challenge — a process that publishing 722 unreviewed manuscripts does not satisfy. Forgoing the million dollars also avoids that debate.

When will the model be accessible?

OpenAI has committed to releasing the model "responsibly," with no date given, and promises workshops as well as future access. The Advisory Group's recommendations (September 29, 2026) push in that direction. But the ecosystem's recent experience — see the Cursor episode — calls for caution about the announced timelines.

Will mathematicians be replaced?

No, but their role is shifting. The model produces results; humans "must still" validate them, understand them, and steer the research — and that is now the bottleneck. The Advisory Group explicitly calls for funding human understanding: value is migrating from demonstration to verification.


✅ Conclusion

In 38 days, an unnamed model produced more mathematical manuscripts than several university departments publish in a year: the question is no longer whether AI can do research, but who will have the means to verify it. Clone the openai/math repository, start with the 162 Lean proofs, and follow the public re-check of the 003 family — that's where, and not in press releases, the credibility of this "new era of discovery" will be decided.