AI model capabilities are advancing extremely rapidly, with benchmark scores and revenue doubling in months.
The case
Anthropic's internal model 2 scores 8.8 percentage points higher on CobBench than Methos Preview, indicating a large gap between internal and publicly released model capabilities.
“On cobbench enthropic enthropics model 2 is about 8.8 percentage points higher than methos preview.”Nathan Labenz · 22 Aug 2026
The median 2026 Stripe founding cohort generates 50% more revenue than the comparable 2025 cohort.
“The median 26 cohort is generating 50% more revenue than the comparable 25 cohort.”Patrick Collison · 17 Aug 2026
Grok 4.7, a significantly more capable model, will be released within a few weeks.
“A larger model, Grock 4.7 that should be significantly more capable is coming out, you know, in a few short weeks.”David Friedberg · 14 Aug 2026
Grok 4.5 uses just one-third the tokens of GPT-5.5 or Fable while achieving a similar intelligence score.
“Grok 4.5 uses just one-third the amount of tokens as GPT-5.5 or Fable while achieving a similar score.”Ryan Greenblatt · 11 Aug 2026
Frontier AI models are causing a massive reduction in the time between vulnerability discovery and exploitation.
“They are causing kind of a massive reduction in the time between the vulnerability discovery and vulnerability exploitation.”Zane Lackey · 7 Aug 2026
The pushback
GPT-4.5 was internally considered a bust by people at OpenAI.
“There's GPT-4.5, which famously people at OpenAI thought was a bit of a bust.”Ryan Greenblatt · 11 Aug 2026
AI model performance advantages evaporate within weeks, as open and closed competitors replicate or exceed frontier capabilities almost immediately after benchmark publication.
“These models are getting commoditized much faster than anybody thought. And how do we know this? Because there is no meaningful sustained advantage once a model publishes their performance criteria. What you see is literally within weeks other models some open some closed some open weight who are able to match and in some cases exceed the performance.”Chamath Palihapitiya · 24 Jul 2026
Fable model showed a regression compared to Opus on subjective nuanced tasks, demonstrating that greater model power does not always improve performance in these domains.
“Fable was actually a big regression, interestingly. which I think is just an interesting point which shows that more power doesn't like when you when you're dealing with these sort of more subjective nuance things, more power doesn't always mean better.”Cameron · 27 Jun 2026
State-of-the-art AI models still fail approximately 20% of the time on fourth-grade science tasks such as boiling water.
“The best models right now are getting something like 80% on the fourth grade science.”Nathan Labenz · 6 Jun 2026
Coding models are approaching a performance plateau, making fine-tuning a viable strategy for use-case-specific optimization.
“We're approaching a certain plateau in how good coding your data to fine-tune a model specifically for your use case.”Amjad Masad · 25 Apr 2026
Reinforcement learning is not ready to learn PCB routing from scratch by interacting with CAD software via keyboard and mouse.
“The reality is that I don't think reinforcement learning as a technology is ready for something.”Sergiy Nesterenko · 15 Apr 2026