# Fireworks AI

> Fast open-model inference plus SFT, DPO and reinforcement fine-tuning at base-model prices.

Canonical: https://www.subconscious.dev/providers/fireworks · By The Subconscious Team · Updated September 30, 2026

- Founded: 2022
- Example models: DeepSeek V4 Pro, Kimi K3
- Website: https://fireworks.ai

## Overview

Fireworks AI was founded in 2022 by former Meta PyTorch engineers led by CEO Lin Qiao, and it sells speed on open models. Its custom serving stack has posted 167 to 174 tokens per second on DeepSeek V4 Pro in third-party measurements, several times what most GPU peers hit on the same model. The catalog holds 400+ models across text, vision, audio and embeddings, served through an OpenAI-compatible API. In July 2026 it raised a $1.505B Series D at a $17.5B valuation, with a reported $1B+ run rate and 40T+ tokens a day.

Post-training is the other half of the pitch. Fireworks offers SFT, DPO and reinforcement fine-tuning in LoRA or full-parameter form, and a fine-tuned model serves at the same per-token price as its base. The Training API went GA on August 31, 2026, letting researchers run their own RL loop on Fireworks-managed trainers and rollout fleets with matched numerics between training and inference. Fireworks says it has run RL across more than 10,000 GPUs, and a Fireworks Lab team embeds with customers on custom builds.

## Upsides

- Among the fastest GPU-based throughput on popular open models.
- No serving markup on fine-tuned models.
- Full 1M context on models like DeepSeek V4 Pro where cheaper hosts truncate it.
- SOC 2, HIPAA and ISO certifications plus AWS and GCP marketplace billing.

## Core use cases

- Latency-sensitive production chat and tool-calling agents.
- Reinforcement fine-tuning an open model to beat a closed API on a narrow task.

## Downsides

- Pricier than bargain hosts like Novita on most smaller commodity models.
- Dedicated GPU rates rose on September 1, 2026, with an H100 now $8 an hour.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open weights |
| Flagship models | DeepSeek V4 Pro, Kimi K3 |
| Speed | 167–174 tok/s on DeepSeek V4 Pro |
| Price | Fine-tunes served at base price |
| Customization | SFT, DPO, RFT; Training API |
| Deployment | Serverless, dedicated GPUs |
| Long context | Full 1M on DeepSeek V4 Pro |

## FAQ

### What is Fireworks AI?

Fireworks AI was founded in 2022 by former Meta PyTorch engineers led by CEO Lin Qiao, and it sells speed on open models. Its custom serving stack has posted 167 to 174 tokens per second on DeepSeek V4 Pro in third-party measurements, several times what most GPU peers hit on the same model. The catalog holds 400+ models across text, vision, audio and embeddings, served through an OpenAI-compatible API. In July 2026 it raised a $1.505B Series D at a $17.5B valuation, with a reported $1B+ run rate and 40T+ tokens a day.

### What is Fireworks AI best for?

Latency-sensitive production chat and tool-calling agents; Reinforcement fine-tuning an open model to beat a closed API on a narrow task.

### How much does Fireworks AI cost?

Fireworks AI pricing at a glance: Fine-tunes served at base price. Rates change often, so check Fireworks AI's pricing page before committing.

### How much context does Fireworks AI support?

Fireworks AI's long-context support: Full 1M on DeepSeek V4 Pro.

### What are the downsides of Fireworks AI?

Pricier than bargain hosts like Novita on most smaller commodity models; Dedicated GPU rates rose on September 1, 2026, with an H100 now $8 an hour.

### What are the best alternatives to Fireworks AI?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Fireworks AI on this site.

## Comparisons

- [Subconscious vs Fireworks AI](https://www.subconscious.dev/compare/subconscious-vs-fireworks.md)
- [OpenAI vs Fireworks AI](https://www.subconscious.dev/compare/openai-vs-fireworks.md)
- [Anthropic vs Fireworks AI](https://www.subconscious.dev/compare/anthropic-vs-fireworks.md)
- [Google Vertex AI vs Fireworks AI](https://www.subconscious.dev/compare/google-vertex-vs-fireworks.md)
- [Amazon Bedrock vs Fireworks AI](https://www.subconscious.dev/compare/aws-bedrock-vs-fireworks.md)
- [Together AI vs Fireworks AI](https://www.subconscious.dev/compare/together-ai-vs-fireworks.md)
- [Fireworks AI vs Baseten](https://www.subconscious.dev/compare/fireworks-vs-baseten.md)
- [Fireworks AI vs Groq](https://www.subconscious.dev/compare/fireworks-vs-groq.md)
- [Fireworks AI vs Cerebras](https://www.subconscious.dev/compare/fireworks-vs-cerebras.md)
- [Fireworks AI vs DeepInfra](https://www.subconscious.dev/compare/fireworks-vs-deepinfra.md)
- [Fireworks AI vs Modal](https://www.subconscious.dev/compare/fireworks-vs-modal.md)
- [Fireworks AI vs xAI](https://www.subconscious.dev/compare/fireworks-vs-xai.md)
- [Fireworks AI vs DeepSeek](https://www.subconscious.dev/compare/fireworks-vs-deepseek.md)
- [Fireworks AI vs Moonshot AI](https://www.subconscious.dev/compare/fireworks-vs-moonshot-ai.md)
- [Fireworks AI vs Z.ai](https://www.subconscious.dev/compare/fireworks-vs-z-ai.md)
- [Fireworks AI vs Alibaba Cloud](https://www.subconscious.dev/compare/fireworks-vs-alibaba-cloud.md)
- [Fireworks AI vs Meta](https://www.subconscious.dev/compare/fireworks-vs-meta.md)
- [Fireworks AI vs SambaNova](https://www.subconscious.dev/compare/fireworks-vs-sambanova.md)
- [Fireworks AI vs Nebius](https://www.subconscious.dev/compare/fireworks-vs-nebius.md)
- [Fireworks AI vs fal](https://www.subconscious.dev/compare/fireworks-vs-fal.md)
- [Fireworks AI vs Novita AI](https://www.subconscious.dev/compare/fireworks-vs-novita-ai.md)
- [Fireworks AI vs Parasail](https://www.subconscious.dev/compare/fireworks-vs-parasail.md)
- [Fireworks AI vs Inference.net](https://www.subconscious.dev/compare/fireworks-vs-inference-net.md)
- [Fireworks AI vs GMI Cloud](https://www.subconscious.dev/compare/fireworks-vs-gmi-cloud.md)
- [Fireworks AI vs Sail Research](https://www.subconscious.dev/compare/fireworks-vs-sail-research.md)
- [Fireworks AI vs Morph](https://www.subconscious.dev/compare/fireworks-vs-morph.md)
- [Fireworks AI vs Relace](https://www.subconscious.dev/compare/fireworks-vs-relace.md)
- [Fireworks AI vs TypeSafe AI](https://www.subconscious.dev/compare/fireworks-vs-typesafe-ai.md)
- [Fireworks AI vs StepFun](https://www.subconscious.dev/compare/fireworks-vs-stepfun.md)
- [Fireworks AI vs Runware](https://www.subconscious.dev/compare/fireworks-vs-runware.md)
- [Fireworks AI vs StreamLake](https://www.subconscious.dev/compare/fireworks-vs-streamlake.md)
- [Fireworks AI vs Wafer](https://www.subconscious.dev/compare/fireworks-vs-wafer.md)
- [Fireworks AI vs RunInfra](https://www.subconscious.dev/compare/fireworks-vs-runinfra.md)
- [Fireworks AI vs Particle.AI](https://www.subconscious.dev/compare/fireworks-vs-particle-ai.md)

## Sources

- [Fireworks Training API GA](https://fireworks.ai/blog/train-past-the-frontier-training-api-now-generally-available)
- [Fireworks review, ChatForest](https://chatforest.com/reviews/fireworks-ai-inference-fine-tuning-platform/)

Pricing and model lineups change often; figures are a snapshot.
