On the final day of August, a human mathematician made a landmark advance on one of number theory’s most enduring problems—only to be overtaken by artificial intelligence within days.
Julia Stadlmann of the University of Illinois Urbana–Champaign announced a new record in the effort to resolve the twin prime conjecture, which proposes that infinitely many pairs of prime numbers differ by exactly two. Examples include 3 and 5, or 17 and 19. Her result marked the first major improvement in the field in more than a decade.
It is a “fiendishly difficult problem,” said Kannan Soundararajan, a mathematician at Stanford University, who praised Stadlmann’s persistence and boldness. “It was really quite impressive.”
Within three days, the AI startup Axiom Math had built on Stadlmann’s methods to surpass her result. Just two hours after Axiom’s announcement, OpenAI went further still, improving the record again. Five days later, the company said it had also solved another major mathematical puzzle.
“It’s a David-versus-Goliath story where the human gets the gold,” said Kevin Ford, Stadlmann’s postdoctoral mentor at Illinois. For three days, at least.
The rapid sequence of breakthroughs has intensified mathematicians’ unease about AI. Many researchers argue that some AI companies are not only disregarding professional norms, but also weakening the field’s broader pursuit of understanding. “What we are seeing here is a consequence of some very deliberate choices to abandon any pretense of gaining human understanding … and using AI tools for the sole purpose of achieving a benchmark,” Terence Tao of UCLA, a Fields Medalist and one of mathematics’ most influential writers, posted on mathstodon.xyz.
Chasing Prime Pairs
Prime numbers—integers divisible only by one and themselves—have long fascinated mathematicians, who continue to investigate how they are distributed along the number line. The twin prime conjecture, studied at least since the 19th century, is widely believed to be true but remains unproven. A more accessible version asks whether there are infinitely many pairs of primes separated by the same gap, even if that gap is larger than two.
In 2013, Yitang Zhang, now at Sun Yat-sen University in Guangzhou, China, proved that the answer is yes: some finite gap must occur infinitely often. Mathematicians did not know whether that gap could be two, but Zhang showed that a gap below 70 million must recur endlessly, establishing the first record for the “small prime-gap bound.”
Researchers soon began lowering that number. Each reduction brings mathematicians closer to proving the twin prime conjecture. By mid-2014, large collaborative efforts had brought the bound down to 246, meaning that infinitely many prime pairs must exist with a gap no larger than 246. That record stood for 12 years.
Stadlmann began working intensively on the problem about two years ago, after completing her doctorate under James Maynard at the University of Oxford in England. Maynard helped establish the previous record and won the Fields Medal in 2022, in part for his work on prime gaps. Stadlmann took on the difficult task of more fully integrating newer techniques with Zhang’s earlier approach.
Julia Stadlmann held the record for bounded gaps between primes for three days. Rod Searcey
Both approaches rely on a “sieve” method, which partially filters out composite numbers using different weights—for example, filtering multiples of two more strongly than multiples of three. Paradoxically, this partial filtering can reveal more about the distribution of primes than a complete filter would. Stadlmann identified optimal weights by assembling regions of high-dimensional space where different methods perform best. She then had to calculate the volumes of those regions with only an incomplete picture of their shapes.
“It goes beyond finding a needle in a haystack,” said Andrew Granville, a number theorist at the University of Montreal.
By early summer 2026, Stadlmann had reduced the world record from 246 to 240. With more time, number theorists say, her methods might have gone even further. But time was precisely what she lacked.
In mid-August, Stadlmann began hearing rumors that OpenAI was preparing to announce a major new result on prime gaps. Ford urged her: “Drop everything you’re doing and put this result on the math preprint archive. Get it out there. Even if it’s improved the next day, you’ve got the record at least for a day.”
As it happened, her record lasted three days. It may be the last time a human holds it.
After Stadlmann posted her result online, researchers at Axiom Math, who had been examining proofs published in 2013 and 2014, redirected their work and incorporated her new ideas into a theorem-proving AI system. Over the next three days, working through the night, the Axiom team lowered the bound to 212—an estimate of what Stadlmann might have achieved with additional time and computing power.
Meanwhile, OpenAI had been working on a closely related problem involving the largest gaps between primes rather than the smallest. The company had used that problem as a benchmark to demonstrate the problem-solving abilities of its newest large language model, GPT-6 Astra. According to OpenAI computer scientist Sébastien Bubeck, the team was also exploring the small-gaps problem as an additional effort because mathematicians at the company were interested in it. Astra reduced the record to 186, OpenAI reported in a paper released alongside the model on September 3.
Bubeck said OpenAI does not plan to push its prime-gap work further. “Our goal is not to preemptively strike and capture all the results possible,” he said. “Our strategy is to empower the mathematician.”
Even so, controversy spread quickly through the mathematics community.
Mathematics and Understanding
Few dispute that AI can be an extraordinarily powerful mathematical tool. “Last week, I did a proof in two hours that would have taken me a month before,” Granville said. AI systems can search vast bodies of literature faster than humans, identify unproductive paths more quickly, and test far more combinations of ideas.
But they also raise difficult questions: Who will have access to these systems? How expensive will they be? How should mathematicians interpret machine-generated proofs that may contain “AI slop” humans cannot immediately understand? And how will the culture of mathematics change if a computer can suddenly make a researcher’s work obsolete?
Although Stadlmann said she does not feel mistreated, many mathematicians were troubled that OpenAI had worked on the same problem without informing her. In mathematics, researchers typically avoid encroaching on problems being pursued by others, especially less senior scholars, or they reach out to collaborate. “I do not see that OpenAI carefully considered how they treated her,” Granville said.
OpenAI said it had been in contact with Maynard but had not reached out directly to Stadlmann.
Mathematicians are now studying OpenAI’s paper and the strategy behind the new record. “I am not so much interested in the exact number, but I am very interested in the mathematical ideas,” Stadlmann said.
That pursuit of ideas is what many mathematicians fear could be lost as AI is unleashed on unsolved problems. For questions like these, Tao noted, the value lies not merely in the final answer, but in what emerges from the struggle to reach it. Along the way, mathematicians develop new techniques, new ways of thinking and a deeper understanding of mathematics as a whole—insights that cannot be captured by a single number such as 186.
“I love the process of problem solving,” Ford said. “Banging our heads against the wall and getting frustrated, and all of a sudden seeing the idea that works.” He worries that as AI takes that process away, talented young people may choose not to enter the field. Future researchers may be discouraged from pursuing problems when a machine can reach the answer first. “That would be massively detrimental,” he said, “not just to math but to science in general.”


