Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Groq

Inference-specialist chipmaker whose LPU delivers record token throughput and low latency.

Founded
2016
Headquarters
Mountain View, California, US
Website
groq.com

Groq is an AI inference company that designs its own Language Processing Unit (LPU) — a deterministic, low-latency chip architecture built specifically for serving models rather than training them. It offers GroqCloud, an API that runs popular open-weight models at very high tokens-per-second, and has positioned speed and predictable inference economics as its core differentiators against GPU-based serving.

Groq was founded in 2016 by Jonathan Ross, who had helped create Google's first Tensor Processing Unit. Its LPU takes a different architectural path from GPUs: a deterministic, software-scheduled design that avoids the variability of GPU inference and excels at the sequential token generation that dominates LLM serving.

GroqCloud serves open-weight models (Llama, Qwen, and others) at throughput and latency figures that are hard to match on general-purpose GPUs, making Groq popular for latency-sensitive and high-volume inference. The company sells both cloud access and on-premises systems, and has pursued large sovereign-AI deployments in the Middle East.

Backed by significant venture and strategic funding at a multi-billion-dollar valuation, Groq is one of the most prominent challengers arguing that inference — not training — is where custom silicon will unseat NVIDIA, competing alongside Cerebras and hyperscaler ASICs for that workload.

Groq sells model access we track in the pricing index. API pricing & models →

Key people

1

Related companies

3

Part of the Tokenando AI Landscape · Explore the landscape → · All companies →