AI benchmark scores and revenues are compounding at rates that make annual planning obsolete
Frontier model scores on the Apex benchmark rose from 1 percent to 40 percent in a single year. Revenue at the leading labs is growing at comparable speed. The pattern across capability and commerce is consistent enough to demand a harder look at what "planning horizon" means in this environment.
Benchmark scores on the Apex evaluation rose from 1 percent to 40 percent in 12 months. Brendan Foody, who tracks frontier model performance, notes that a year ago the leading model on Apex was scoring at the 1 percent level. The frontier now sits at 40 percent. That is not incremental progress on a familiar curve. It is a step change in what models can attempt, compressed into a single calendar year.
The revenue picture matches. Krishna Rao describes Anthropic’s annualized run-rate growing from roughly $9 billion at the start of the year to north of $30 billion by the end of one quarter. Chamath Palihapitiya puts Anthropic’s ARR at over $70 billion, against an internal forecast targeting $100 billion for the full year. OpenAI’s trajectory is similarly steep: Palihapitiya reports the company’s ARR rose from $33 billion in May to $41.3 billion in July, with the year-end forecast revised upward from $60 billion to approximately $75 billion. Marc Andreessen argues that Anthropic and OpenAI are adding more revenue per month than Meta, Google, or Microsoft, and puts the combined run-rate of the two companies on a path to $200 billion by year-end. Harry Stebbings adds that Anthropic’s token volume grew approximately 15 times in the first quarter alone.
The speed of model improvement inside the capability curve is equally steep. Dylan Patel observes that in two months, model capabilities advanced from L4 to L6 engineer level by his reckoning. Nathan Labenz, drawing on METR’s task-length analysis, notes a doubling time for AI task length of under four months, implying an 8 to 12 times annual gain. Mark Chen says scaling laws have held across almost 10 orders of magnitude and sees no reason they should stop. Rao adds, from Anthropic’s vantage point, that the scaling laws are not slowing down.
We started the year with about $9 billion of run rate revenue and we ended the quarter with, you know, north of $30 billion of run rate revenue. Krishna Rao
The security domain provides one of the most concrete illustrations of what that trajectory means in practice. Rao describes an open-source codebase where a prior model found 22 security vulnerabilities. Anthropic’s Mythos model found 250. Zane Lackey frames the practical consequence directly: these systems are causing a massive reduction in the time between vulnerability discovery and exploitation. The gap between what an automated audit can catch and what defenders assume it can catch is widening faster than most security teams have updated their baseline assumptions.
The competitive tempo at the frontier is itself an unusual phenomenon. David Sacks describes leapfrogging between frontier models occurring at a cadence of two to four weeks. Martin Casado observes that any given model remains relevant for only three to nine months before being superseded. Gavriel Cohen, who builds agent systems, draws out the operational implication: unlike conventional enterprise software, which can be deployed and left running for years, agents require constant model upgrades, and every upgrade changes the behavior of what has been built. The substrate is in motion underneath every deployment.
The revenue acceleration visible in the two largest labs also appears at the product layer. Amjad Masad reports that Replit grew from $2.5 million to $250 million in annual revenue in one year, and is on a path to $1 billion in the current year. David Haber describes a company whose successive $100 million ARR milestones took 20 months, then 10 months, then 5 months, each interval half the previous one. Patrick Collison notes that Stripe’s internal Minions agent system grew from 1,200 pull requests per week to 7,000 in the period since January or February, and that the median 2026 founding cohort on Stripe’s platform is generating 50 percent more revenue than the comparable 2025 cohort. Chris Degnan captures the competitive pressure this creates: going from $50 million to $100 million in a year, which would have been celebrated five years ago, now risks being insufficient.
What the evidence taken together describes is not a single record-breaking quarter or an anomalous benchmark result. It is a pattern in which the rate of change is itself accelerating. The doubling times for task capability, the halving times for successive revenue milestones, the shrinking window of model relevance, and the leapfrog cadence at the frontier are all pointing in the same direction at once. Whether that trajectory holds at its current slope or moderates, the institutions and companies organized around the assumption that AI progress moves at a conventional pace are already operating with outdated maps.