Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Fireworks AI

Fast, production-grade inference platform for open-weight models and compound AI systems.

Founded
2022
Headquarters
Redwood City, California, US

Fireworks AI is an inference platform that serves open-weight models at high speed and low cost, aimed at production workloads and agentic "compound AI" systems. Founded by engineers from Meta's PyTorch team, it emphasizes serving efficiency, fine-tuning, and a developer experience geared toward shipping AI products rather than experimentation.

Fireworks AI was founded in 2022 by former Meta engineers who had worked on the PyTorch framework and Meta's large-scale inference systems. That heritage informs its focus on squeezing maximum throughput and minimum latency out of open-weight model serving.

The platform offers an OpenAI-compatible API across a wide range of open models, plus fine-tuning, function calling, and tooling for building multi-model agentic systems. Its performance engineering — custom kernels and serving optimizations — targets teams that have outgrown generic inference and need predictable production economics.

Backed by leading venture firms at a multi-billion-dollar valuation, Fireworks competes with Together AI, Groq, Baseten, and hyperscaler serving, riding the same wave of improving open-weight models that makes self-hosted inference an attractive alternative to closed frontier APIs.

Fireworks AI sells model access we track in the pricing index. API pricing & models →

Related companies

4

Part of the Tokenando AI Landscape · Explore the landscape → · All companies →