Groq
Inference-specialist chipmaker whose LPU delivers record token throughput and low latency.
Groq is an AI inference company that designs its own Language Processing Unit (LPU) — a deterministic, low-latency chip architecture built specifically for serving models rather than training them. It offers GroqCloud, an API that runs popular open-weight models at very high tokens-per-second, and has positioned speed and predictable inference economics as its core differentiators against GPU-based serving.
Groq was founded in 2016 by Jonathan Ross, who had helped create Google's first Tensor Processing Unit. Its LPU takes a different architectural path from GPUs: a deterministic, software-scheduled design that avoids the variability of GPU inference and excels at the sequential token generation that dominates LLM serving.
GroqCloud serves open-weight models (Llama, Qwen, and others) at throughput and latency figures that are hard to match on general-purpose GPUs, making Groq popular for latency-sensitive and high-volume inference. The company sells both cloud access and on-premises systems, and has pursued large sovereign-AI deployments in the Middle East.
Backed by significant venture and strategic funding at a multi-billion-dollar valuation, Groq is one of the most prominent challengers arguing that inference — not training — is where custom silicon will unseat NVIDIA, competing alongside Cerebras and hyperscaler ASICs for that workload.
Groq sells model access we track in the pricing index. API pricing & models →
Key people
1Related companies
3Part of the Tokenando AI Landscape · Explore the landscape → · All companies →