What is Groq?
Groq is an AI chip company whose Language Processing Units (LPUs) are designed for deterministic latency, a feature the company has advertised.
Company timeline
- Mar 2026 - Stefano Ermon contrasted his software-based solution with specialized AI inference chips like Groq’s, noting that GPUs are more readily scalable.
- Apr 2026 - Philip Kiely suggested that 2026 could be the year of disaggregation, with Nvidia potentially buying Groq to handle specialized pre-fill and decode compute.
- May 2026 - Reiner Pope noted that Groq has advertised deterministic latency in its processors, but achieving both deterministic latency and high speed simultaneously is challenging.
- Sep 2026 - Sean Lie stated that running a frontier-level model of a few trillion parameters would require thousands of Groq LPUs just to hold the weights.
Where it appears in the record
Every line below is attributed to a named speaker.
CPUs with deterministic latency are technically feasible but have been abandoned because the market does not reward that design trade-off.
“You can actually design a CPU that has deterministic latency as well. In fact, the processors inside a lot of AI chips also have deterministic latency. Groq has advertised this. TPUs have that in the core as well. The challenge is getting deterministic latency and high speed at the same time. Non-deterministic latency comes from specific design choices in a CPU. It's actually possible to remove those design choices and make a CPU with deterministic latency, but those are not very attractive in the market, so people don't make those CPUs anymore.”Reiner Pope · 22 May 2026
Running a frontier model of a few trillion parameters on Groq LPUs would require thousands of chips just to hold model weights, due to insufficient on-chip SRAM, constraining Groq to smaller models.
“If you think about it to run a frontier level model of like let's say a few trillion parameters you need thousands and thousands of Grock LPUs just to hold the weights.”Sean Lie · 2 Sep 2026
Groq, acquired by Nvidia, is positioned as specialized pre-fill compute hardware central to 2026 inference disaggregation trends.
“One thing 2026 could be is the year of disaggregation with, for example, Nvidia buying Groq and having more specialized pre-fill compute versus decode compute.”Philip Kiely · 30 Apr 2026
Software-based diffusion LLM inference on commodity GPUs is more scalable than specialized AI inference chips from Cerebras and Groq, according to Stefano Ermon.
“Our solution is software based so it's much more scalable we're still running on GPUs. So you can you know get as much capacity as you as you know as you can get GPUs for which is relatively easier compared to specialized AI inference chips.”Stefano Ermon · 26 Mar 2026