Fireworks AI

Fast, production-grade inference platform for open-weight models and compound AI systems.

Founded
2022
Headquarters
Redwood City, California, US

Fireworks AI is an inference platform that serves open-weight models at high speed and low cost, aimed at production workloads and agentic "compound AI" systems. Founded by engineers from Meta's PyTorch team, it emphasizes serving efficiency, fine-tuning, and a developer experience geared toward shipping AI products rather than experimentation.

Fireworks AI was founded in 2022 by former Meta engineers who had worked on the PyTorch framework and Meta's large-scale inference systems. That heritage informs its focus on squeezing maximum throughput and minimum latency out of open-weight model serving.

The platform offers an OpenAI-compatible API across a wide range of open models, plus fine-tuning, function calling, and tooling for building multi-model agentic systems. Its performance engineering — custom kernels and serving optimizations — targets teams that have outgrown generic inference and need predictable production economics.

Backed by leading venture firms at a multi-billion-dollar valuation, Fireworks competes with Together AI, Groq, Baseten, and hyperscaler serving, riding the same wave of improving open-weight models that makes self-hosted inference an attractive alternative to closed frontier APIs.

Fireworks AI sells model access we track and price. API pricing & models →

Related companies

4

Part of the Tokenando AI Landscape · Explore the landscape → · All companies →