AI is resolving problems that defeated human researchers for decades, and the pattern spans every field
From quantum mechanics to formal mathematics to protein design, AI systems are closing out problems that resisted solution for 30, 80, and 130 years. The results are concrete enough, and consistent enough across domains, that they require a harder accounting of what scientific progress will look like from here.
Codex solved a simulation of the SYK model, a technically demanding problem in quantum mechanics and gravity, in 10 minutes. Alex Lupsasca, a physicist who described the episode, noted that multiple research groups had been trying to run that simulation and had not managed it. The model did. That is one data point. The accumulation of them is what demands attention.
Lupsasca also described a separate result in particle physics. A formula for single-minus gluon amplitudes, which human researchers had been working on for a year, was first conjectured by a model he referred to as “GPT 5.2 pro,” then proved by an internal OpenAI model. The same model, anchored on the gluon paper, independently completed the graviton amplitude calculation, which Lupsasca described as mathematically quite different from the gluon problem. Brian Greene, separately, says that a tool reproduced months of his string theory results within half an hour, and he sees “a real possibility of the nature of research changing” over the next five to 10 years. Greene also relays a claim he described with some uncertainty: “An 80-year-old mathematical conjecture. I believe it’s called the Jacobian conjecture was solved very recently by a mathematician using AI.” The qualifier is his own.
In mathematics, Carina Hong reports results from Axiom Math that are equally concrete. The system scored 120 out of 120 on the 2025 Putnam exam. The best human score, from a student at either MIT or the University of Chicago (the precise school is not announced), was 110. The best large language model, DeepSeek, reached 103. Beyond the benchmark, Hong describes a team member who used AI methods to disprove a 30-year-old conjecture by finding a counterexample and to find the solution to what she called a 130-year-old problem. Separately, auto-formalization caught something that 50 years of human teaching had not: an implicit assumption in the “agree to disagree” theorem by Nobel Prize winner Robert Aumann, present since 1976 and never made explicit, was identified by a prover during the formalization process and patched.
Thanks to LLM coding, what took a whole team of research scientists at DeepMind and millions of dollars of research and compute can now be done for a few thousand dollars of rented compute. Eric Jang
Biology and materials science show the same pattern at different scales. David Sinclair describes work in his laboratory that would have taken 160 years and, in his words, “quite literally billions of dollars,” now running on a $10,000 budget. Matt McPartlon, whose work involves antibody design, reports that wet-lab validation cycles have dropped from years to weeks. His team designed antibodies to 50 targets and got binders to about half, with an average hit rate of around 20 percent for binding. Design models, he notes, can achieve hit rates above 50 percent, meaning a standard 96-well plate yields roughly 48 viable binders. Nathan Labenz relays that Anthropic reported Claude designed working protein binders. Bo Wang describes Xaira’s X-Cell model predicting how unseen cell lines respond to genetic perturbations, including T-cells the model had never been trained on, accurately across thousands of genes and thousands of perturbations. Joseph Krause reports that his organization’s AI scientist has moved into elemental and alloy families no human researcher has ever published on.
The cost compression Eric Jang describes for AI research itself applies equally across these domains. What required a whole team of research scientists at DeepMind and millions of dollars of compute, he says, can now be done for a few thousand dollars of rented compute. Dylan Patel puts a comparable figure on economics research: a project that would have required a team of 200 economists working for a year was completed by a single person using AI coding tools.
Mark Chen, who works at OpenAI, observes that at AI labs, research work is shifting toward orchestration: researchers generate ideas, and models handle implementation and execution. That structural change in how research is conducted is at least as significant as any individual result. When the labor of implementation no longer scales with the complexity of the problem, the bottleneck moves entirely to the quality of the question being asked.
What the evidence describes is not acceleration along a familiar curve. It is a change in which problems are tractable and for whom. Problems that resisted solution because they required resources, headcount, or time that no single group could sustain are now, case by case, being resolved. The institutions, funding structures, and career paths built around the old cost structure were shaped by constraints that are visibly weakening. The pattern across mathematics, physics, biology, and materials science is consistent enough that treating any one result as an outlier is no longer defensible.