1 Sep 2026
Citation Bureau
Vol. I
No. 300

Custom fine-tuned models deliver significantly lower latency and cost than frontier models while maintaining equal or higher quality.

The case

A 10-to-1 price discrimination ratio from frontier model providers makes it effectively impossible for app-layer companies to compete while using those providers' APIs.

“At a 10 to1 price discrimination ratio it's just really hard for them to compete.”
Nathan Labenz · 22 Aug 2026

Some capabilities in human behavior modeling cannot be achieved through prompting alone; model parameters must be modified.

“There are certain things you just cannot shape just by prompting the model. So some to some degree you do need to touch the parameters of the model itself.”
Joon Sung Park · 21 Aug 2026

The open-source model GLM2 is 90% cheaper than Claude Opus.

“It is 90% cheaper to use these products.”
David Friedberg · 14 Aug 2026

Grok 4.5 uses just one-third the tokens of GPT-5.5 or Fable while achieving a similar intelligence score.

“Grok 4.5 uses just one-third the amount of tokens as GPT-5.5 or Fable while achieving a similar score.”
Ryan Greenblatt · 11 Aug 2026

Silico's $1,000/month price will decrease over time as Goodfire develops more token-efficient agent methods.

“My hope is we're starting out with $1,000 a month subscription. We'll be able to bring that down over time because we're able to come up with more and more clever ways.”
Dan Balsam · 8 Aug 2026

Open-weight models can reach speeds of up to ~500 tokens per second, which is 2-3x faster than the proprietary fast mode available today.

“For proprietary model there is regular mode and fast mode and that's only the two switch here. But for openweight when you're running it, every provider can offer potentially even 10 different levels of speed going from like the slowest mode which can be a lot cheaper to 400 tokens per second almost up to 500 in many cases that for some workloads and this is typically 2x or 3x faster than the fast mode out there today.”
Simon Mo · 6 Aug 2026

The pushback

Ben Thompson's cost analysis found that Kimi K3 does not have a significant cost advantage over US frontier models, countering the narrative of Chinese open-source models being dramatically cheaper to run.

“Ben Thompson on his blog went through some of the cost numbers and it turns out that Kimmy K3 is not that much cheaper to run. There's not a significant cost advantage to it.”
David Sacks · 24 Jul 2026

Frontier labs are capturing an increasing share of economic value (wallet share) even as commodity token volume grows and shifts to cheaper models.

“The share of wallet is actually increasing to the Frontier Labs while the share of tokens, these commodity tokens is obviously going up.”
Brad Gerstner · 11 Jul 2026

Anthropic has become a premium product since the end of last year, priced at twice the cost of its competitor.

“It has become a premium product since the end of last year. It is twice the cost of its competitor, right?”
Jason Lemkin · 28 May 2026

A vulnerability discovered approximately two weeks prior in LiteLLM was found to steal all user keys and credentials, making it unsafe for enterprise deployment.

“We have seen with light LLM, how long is this now ago, 2 weeks or something? You probably saw it, right? Like with this vulnerability that still all of a sudden steals all your keys and credentials and so on and so forth.”
Philipp Herzig · 23 Apr 2026

Topics

AI Cost EfficiencyAI ModelsFine-Tuning

Citation Bureau · compiled from attributed public discussion. Last updated 2026-08-22.