In what is quickly becoming an explosive controversy in mathematics, a potential solution to one of the field’s greatest open questions has been overshadowed by allegations of espionage and unethical academic work.
Over the previous 48 hours (as on September 9 morning IST), the mathematical community has been rocked by claims that OpenAI leveraged its enormous computing power — and potentially the private user data of academic researchers — to scoop a specific solution to the Navier-Stokes equations problem.
The incident has already triggered a debate about the future of human mathematicians in the age of advanced large language models (LLMs), with New York University mathematician Tristan Buckmaster describing the event as a “Deep Blue-Kasparov moment” for the field.
The mathematical problem
The magnitude of the controversy owes itself to the significance of the mathematical problem. The Navier-Stokes equations describe the motion of viscous fluids like water and air, and they have frustrated mathematicians and physicists for more than two centuries.
As OpenAI wrote on its website: “The Navier–Stokes equations use Newton’s second law of motion (’force equals mass times acceleration’) to describe how fluids move. Importantly, they treat a fluid as a continuous medium rather than tracking individual molecules. These equations are used for aircraft design, weather forecasting, and the study of blood flow.”
The Clay Mathematics Institute in New York designated their resolution as one of the seven Millennium Prize Problems. This means anyone who can prove whether the equations always have smooth, predictable solutions or whether the equations can ‘blow up’ — meaning the mathematical solutions reach infinity in finite time — will win $1 million.
The official Millennium Prize problem statement, authored by U.S. mathematician Charles Fefferman, offers multiple avenues for a valid proof. Options A and B deal with unforced equations while Options C and D allow for a solution that includes an external forcing function.
Crucially, for a forced blow-up to qualify for the Millennium Prize under Options C and D, the external force must be spatially and temporally smooth. Put another way, the force causing the blow-up cannot suddenly appear or disappear or change sharply from one place or time to another. It must behave in a predictable way.
For years, the Spanish mathematicians Diego Córdoba and Luis Martínez-Zoroa have championed a specific programmatic approach to solving this problem. Their strategy depends on building blow-ups by repeatedly making small but concentrated changes to a fluid’s motion. Each change induces fast and fine variations, which are added to the existing solution step by step. The aim is for these small ‘corrections’ to pile up and eventually cause the fluid’s behaviour to become infinite or undefined in finite time.
This way, the duo showed that slower, larger-scale parts of the fluid can strengthen smaller, faster-moving waves until the equations develop a singularity (i.e. become infinite or undefined). This means the fluid’s behaviour breaks down in a mathematical sense. Their work has demonstrated the existence of such blow-ups, albeit only when the external force is irregular rather than smooth.
Collaborating in what they have said is a strictly personal capacity, Buckmaster and Anthropic researcher Levent Alpöge have spent the last year using LLMs to push the Córdoba and Martínez-Zoroa programme to completion.
Specifically, they have sought to prove that blow-ups can occur even when the external force is smooth and well-behaved. By mid-August, they had reportedly achieved this for the incompressible porous media equation, the 2D Boussinesq equation, and, notably, the 3D incompressible Euler equations.
They had also made substantial progress on a modified form of the Navier–Stokes equations called the hypo-dissipative equations. These results suggested that their way could be capable of producing blow-ups under increasingly realistic mathematical conditions, bringing them closer to the much harder problem of proving a blow-up for the standard 3D Navier–Stokes equations.
This route — proving finite-time blow-up via smooth forcing using Fefferman’s Options C and D — is also a niche vector of attack. Mathematicians agree that it is not a direction one can arrive at in a few days by simply feeding the Millennium Prize problem into an AI chatbot. This detail is relevant to what followed.
Suspicious events
The events of the last month suggest an escalating arms race between academic deliberation and Big Tech’s brute-force computational methods.
On August 15, after months of slow progress, Buckmaster and Alpöge achieved their breakthrough, obtaining blow-up results with smooth forcing for both the Boussinesq and the Euler approximations.
A week later, the duo had their proofs — heavily generated by models like Anthropic’s Claude and OpenAI’s GPT-5.6 Sol — checked by a computer to make sure the mathematical arguments were correct. However, the raw output from the AI was still practically unreadable: Buckmaster described the initial LLM-generated Euler proof as “the most horrendous I have ever read” and called it “AI slop”.
So instead of rushing to publish, Alpöge and Buckmaster spent weeks working around the clock to translate the AI output into a readable, traditional mathematical paper.
But then, in early September, a rumor began circulating within the tech industry that Anthropic had solved a major open problem in mathematics. On September 3, after hearing that OpenAI was investigating the rumors, Buckmaster preemptively emailed a prominent mathematician at OpenAI.
This information comes from a statement that Buckmaster published on September 8/9 on his New York University webpage.
Buckmaster said he clarified that his work with Alpöge was in a personal capacity, that it did not represent an institutional effort by Anthropic, and that the duo planned to post its results once it had finished polishing the manuscript. The (then unnamed) interlocutor at OpenAI replied on the same day, expressing excitement and offering computational resources.
But on September 4, OpenAI abruptly requested a same-day meeting with Buckmaster, which Buckmaster deferred to the following week (i.e. the current week). Two days later, OpenAI urgently requested a meeting “at any point today”. At this point, Buckmaster said he joined a call with Sebastien Bubeck and another OpenAI representative. Bubeck is a computer scientist at OpenAI. Alpöge did not attend the call.
During this call, Buckmaster said Bubeck dropped a bombshell: an internal OpenAI model had produced a 100-page proof of finite-time blow-up for the forced Navier-Stokes equations. When Buckmaster pressed for details, the OpenAI team revealed that their solution was exactly: “Existence of forced blowup in R3 and T3, with the forcing function is smooth Option C and D in Fefferman.”
Buckmaster wrote that he immediately recognised this as a “bright red flag” because it was the exact mathematical route he and Alpöge had been pursuing. Significantly, Buckmaster wrote that he and Alpöge had been logging their work on OpenAI Codex — a tool to translate natural language prompts into code — for months.
At this point, Buckmaster refused to cooperate with OpenAI and rushed to publish his and Alpöge’s results, accompanied by his public statement on the confrontation.
OpenAI subsequently held a press briefing in which it claimed 10,000 AI agents had cracked the problem in an 88-hour marathon that had cost the company millions of dollars.
Allegations
Buckmaster’s public statement has already triggered intense scrutiny over OpenAI’s corporate practices, revolving around two major allegations: data privacy violations and academic credit extortion.
When OpenAI claimed to have solved the precise formulation of the problem that Buckmaster and Alpöge had also been working on, Buckmaster questioned how the AI model had arrived at that solution. During the September 6 call, Bubeck had allegedly claimed at first that the model was given “very little human input” and that it had simply been prompted with the Millennium Prize problem statement.
However, as the call progressed, this narrative allegedly unravelled. According to Buckmaster, Bubeck later admitted that an entire team had been assigned to crack the problem and that team members had set the model on easier equations like the Euler approximation first, all while spending an “insane amount of compute”.
Of greater concern: Buckmaster said he had been inputting all of his and Alpöge’s project drafts into OpenAI Codex over the last year. So when Buckmaster asked point-blank if the internal OpenAI model had been trained on or had access to his private Codex sessions, OpenAI stated that the model “did not look up user data”. But when Buckmaster pressed specifically on whether his data had been used to train the model that worked on the problem, OpenAI didn’t answer.
To many in the tech and maths communities, this non-answer stood to mean OpenAI may have effectively ingested an academic’s unpublished work, repackaged it using the help of large server farms, and prepared to claim the glory of cracking a Millennium Prize problem.
But perhaps the most damning allegation concerns OpenAI’s attempt to manipulate the authorship of the discovery. According to Buckmaster, Bubeck offered him two manifestly unethical proposals. The first was that Buckmaster would post his Euler result and that OpenAI would post their Navier-Stokes result the following day.
The second proposal, if Buckmaster’s statement is fully faithful to reality, was even more murky: that Buckmaster write a solo paper — i.e. excluding Alpöge — presenting the Navier-Stokes result, and acknowledge the internal OpenAI model. Bubeck reportedly asserted twice during the call that he wanted Alpöge completely removed from authorship because Alpöge is an employee of Anthropic, OpenAI’s chief rival. OpenAI allegedly sweetened the deal by promising that if the group agreed to the second proposal, the company would publicly state that Buckmaster deserved the $1 million Clay Prize and would acknowledge him to be the “closest human to the problem”.
In essence, OpenAI was allegedly offering an academic a highly prestigious achievement in mathematics on a platter provided the academic erased his collaborator from the record to deny a rival company the PR success it craved.
Buckmaster said he refused both offers and threatened to go public with the conversation, whereupon Bubeck allegedly responded: “Why would you ruin your career?” and “If you don’t want me to be nice, then I don’t have to be nice”. Later that night, Alpöge reportedly also received a text message suggesting Buckmaster was not being “fully rational”, in an apparent attempt to drive a wedge between the two collaborators.
Bubeck has since publicly called the allegations against him “false and inflammatory”.
Nature of mathematics
Beyond the corporate drama, the Buckmaster-Alpöge papers highlight a paradigm shift in how mathematical research has been happening. The duo did not prove their results using the traditional pen-and-paper (or for that matter chalk-and-chalkboard). Instead, they used LLMs to generate the proofs, acting more as conductors than the traditional authors.
However, Buckmaster’s statement also comes with a warning about the quality of AI-generated mathematics. He lamented the “presentation quality” of their rushed release and noted that the community deserves papers written with deep human care.
Like the other sciences, mathematics relies on comprehension and human intuition. But if AI models can generate 100-page proofs in 88 hours, except those proofs are unreadable “slop” that only another machine — like a computer — can verify, the human element of mathematics fades.
As Buckmaster noted, he originally planned to announce his work by emphasising the reality that a mathematician and an LLM can now do a year’s worth of elite mathematics research in one month.
The speed has also proved a double-edged sword now, however: the conventional academic process involves spending weeks or months writing up a verified result so that others (i.e. humans) can learn from it. But OpenAI’s actions suggest that the technology behemoths behind the AI models perceive this window of carefulness to really be a liability.
Specifically, if an academic takes time to polish a preprint paper, a tech giant with millions of dollars in compute can catch wind of the result, brute-force the remaining logic, and publish first.
Unanswered questions
As the dust continues to not settle on the controversy, there remain several open questions.
1. Did OpenAI truly solve the Navier-Stokes equations problem? Because while OpenAI has said its agents reached a solution in 88 hours, it is unclear who, if any, outside the company has seen the claimed 100-page proof. Further, Fefferman’s own problem statement is famously strict. Since Buckmaster described the AI-generated paper as “slop”, it remains to be seen whether OpenAI’s internal model actually produced a mathematically rigorous proof or merely a convoluted approximation that is inaccessible to humans.
Addendum: What value accrues to OpenAI and the AI models it is purveying to claim that thousands of its AI agents solved a tough maths problem within four days at the cost of millions of dollars?
2. Did OpenAI breach user privacy? The silence regarding whether OpenAI trained its latest frontier model on Buckmaster’s private Codex logs is deafening. If OpenAI routinely scrapes the private sessions of enterprise or academic users to identify and solve the exact problems those users are working on, it represents a dire breach of trust that could push academics to abandon cloud-based LLMs altogether.
3. Who deserves the credit? That is, if OpenAI’s proof is valid, who gets the Clay Millennium Prize? Is it the AI agents? Is it OpenAI as a corporation? Is it Buckmaster and Alpöge, whose methodology the AI seemingly copied? Or is it Diego Córdoba and Luis Martínez-Zoroa, who invented the fundamental programme that made the LLM’s brute-force attack possible in the first place? In fact, Buckmaster wrote in his statement: “Let me make plain what I have said to colleagues in private: in view of this body of work, I believe Luis Martínez-Zoroa deserves a Fields Medal.”
Aside 1: A logical extension to this could practically encompass all of humankind, considering the companies behind these powerful AI models trained those models on knowledge generated by millions upon millions of people on the internet — which the companies have never paid for.
Aside 2: As Grace Huckins wrote on MIT Tech Review, “there might be a thin silver lining to that version of the story for mathematicians, because it would suggest that the hard work of two humans, one of whom is a prominent expert on Navier-Stokes, was essential to the agents’ ability to solve the Millennium Problem. Experts have long identified “research taste,” or the ability to choose promising research questions and directions, as a major obstacle for AI in science and mathematics.

4. Where does mathematics go from here? Buckmaster assessed that the solution to the problem represents a “Deep Blue-Kasparov moment”. When the IBM Deep Blue computer defeated Garry Kasparov, chess did not die but the nature of the game changed forever. At the least, organisers had to institute strict bans on computer assistance to maintain the integrity of competitions.
Of course, mathematics is not a competition — at least in its original conception as a human enterprise. If the ‘discovery’ of new mathematical truths is outsourced to black-box corporate server farms that operate faster than humans ever could, the academic community must decide how to train students and assign credit in a way that keeps the field interesting to humans.
As Huckins put it, “Over the past few months, I’ve heard from several researchers that mathematicians are becoming depressed, and it’s not difficult to see why. Mathematics is quickly becoming the province of frontier AI companies with impressive internal-only models, money to burn, and a lack of collaborative spirit.”
In fact, just before the controversy broke out, Australian-American mathematician Terence Tao had posted to his fediverse account a series of notes about why a mathematics by and for humans remains relevant in the age of AI. Among other things, he wrote, “The point remains that there is a substantial opportunity cost in converting a historically productive and motivating problem (such as Navier-Stokes regularity) into a mere viral social media post advertising some benchmark progress, rather than actually advancing the field and developing the next generation of both problems to ask, and people to work on them.”