AI is solving mathematical and scientific problems that defeated human researchers for decades
From a perfect Putnam score to a simulation that stumped research groups for over a year, AI systems are closing cases that were considered intractable. The results span mathematics, physics, and biology, and the cost compression accompanying them is as significant as the breakthroughs themselves.
Sean Carroll, the physicist and author, puts it plainly: an OpenAI model recently solved the famous Erdős unit distance problem. Greg Brockman, OpenAI’s co-founder, states that the company has significant progress on another Millennium Prize problem. That is a hedged claim, and it should be read as one: Brockman’s own words are that OpenAI “have significant progress on another one of these millennia problems,” without specifying which problem or how close a solution is. These are not closed cases. They are indicators of a trajectory.
The mathematics evidence reaches well beyond those claims. Brian Greene, the physicist, reports that an 80-year-old conjecture known as the Jacobian conjecture was recently solved by a mathematician using AI. Carina Hong, whose work at Axiom Math focuses on automated formal proof, reports that her team’s system disproved a 30-year-old conjecture by finding a counterexample and found the solution to a 130-year-old problem involving the global Leono function. On the 2025 Putnam exam, Axiom Math’s system scored a perfect 120 out of 120 points, against a best human score of 110 and a best large language model score of 103 from DeepSeek. Hong also reports that auto-formalization caught and patched an implicit assumption in a 50-year-old Nobel Prize-winning theorem in economics, attributed to Robert Aumann, that had never been made explicit in half a century of teaching.
In physics, Alex Lupsasca, a theoretical physicist, describes receiving word that Codex wrote a simulation of the SYK model, a technical problem in quantum mechanics and gravity, in 10 minutes after multiple research groups had been unable to complete it. He also describes a case in which, using ChatGPT, a physics problem that had been puzzling researchers was solved before a visiting senior researcher even arrived. On the problem of single-minus gluon amplitudes, Lupsasca reports that “the final formula was first conjectured by GPT 5.2 pro and then proved by an internal OpenAI model.” A formula that human researchers had been unable to simplify, conjectured and proved by two AI systems in sequence.
The final formula was first conjectured by GPT 5.2 pro and then proved by an internal OpenAI model. Alex Lupsasca
The cost compression accompanying these results is as striking as the results themselves. David Sinclair, the biologist, describes work in his laboratory that would have taken 160 years and, in his words, “quite literally billions of dollars” now running on a $10,000 budget. Eric Jang states directly that what once required a whole team of research scientists at DeepMind and millions of dollars of compute can now be done for a few thousand dollars of rented compute. Dylan Patel describes a research project completed by a single person using Claude Code that would, by his estimate, have required a team of 200 economists working for a year.
Biology is following a similar curve. Matt McPartlon, whose work involves computational antibody design at Chai Discovery, reports that protein design models can now achieve hit rates above 50 percent, and that the wet-lab validation cycle for designed proteins has dropped from years to weeks. Alex Rives, whose work involves protein language models at EvolutionaryScale, reports that searching the ESMC model can find antibodies reaching the level of affinity required for therapeutic function. The first version of the ESM Atlas was used to find a new gene editing system. Jeff Coller describes AI trained on millions of viral sequences designing novel bacteriophages never seen in nature, with 16 of them outperforming naturally occurring viruses.
Terence Tao, one of the world’s leading mathematicians, offers a quieter data point: people used the latest AI tools to gather numerical evidence in the course of mathematical work. That detail is less dramatic than a Millennium Prize claim, but in some ways more telling. When working mathematicians at that level treat AI as a standard part of the research apparatus, the tool has crossed a threshold that hype alone cannot explain.
Mark Chen, a researcher at OpenAI, describes the internal shift that follows from all of this: at AI labs including OpenAI, he says, the work is becoming mostly orchestration-focused, with researchers generating ideas and models handling implementation and execution. That reorientation is not a prediction. It is a description of what is already happening inside the institutions closest to these tools. The question for every other institution is how long before the gap between what AI can now attempt and what they are organized to do becomes impossible to manage.