# Subconscious vs Fireworks AI

> Fireworks is fast on open-model calls. Subconscious is fast where agents spend their time, on long traces past 200K tokens, and bills only the tokens it processes.

Canonical: https://www.subconscious.dev/compare/subconscious-vs-fireworks · By The Subconscious Team · Updated September 30, 2026

## How they compare

Fireworks and Subconscious both chase speed on open weights, but they measure it in different places. Fireworks has posted 167 to 174 tokens per second on DeepSeek V4 Pro in third-party tests, and it serves the full 1M context on that model where cheaper hosts truncate. That is raw decode speed on a tuned serving stack. Subconscious attacks the growing context instead. Its runtime prunes the KV cache and preserves suffix state, and it delivers 2x faster task completion, 50% to 80% lower cost than standard inference and a 5M+ effective context window. On a trace that sends 1M tokens, Fireworks bills for 1M. Subconscious bills for what survives compression, which might be 200K.

Fireworks wins on post-training breadth. SFT, DPO and reinforcement fine-tuning, a Training API with matched numerics, fine-tunes served at base price and a 400+ model catalog make it the better choice for tuning an open model to beat a closed API on a narrow task, or for latency-sensitive chat. Subconscious is the better fit when the agent is the product and the trace keeps growing: hour-long coding sessions, research agents and review pipelines, where it delivers neutral to 10% better scores on agentic benchmarks.

## What each one does

### Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### Fireworks AI

Fireworks AI was founded in 2022 by former Meta PyTorch engineers led by CEO Lin Qiao, and it sells speed on open models. Its custom serving stack has posted 167 to 174 tokens per second on DeepSeek V4 Pro in third-party measurements, several times what most GPU peers hit on the same model. The catalog holds 400+ models across text, vision, audio and embeddings, served through an OpenAI-compatible API. In July 2026 it raised a $1.505B Series D at a $17.5B valuation, with a reported $1B+ run rate and 40T+ tokens a day.

## Which is best, and when

### Choose Subconscious for

- Traces that send 1M tokens but should not bill for all of them
- Hour-long coding agents past 200K tokens
- Long-horizon agents needing more than 1M tokens of context

### Choose Fireworks AI for

- Reinforcement fine-tuning an open model for a narrow task
- Teams that want a 400+ model catalog on one API
- Latency-sensitive chat and tool calls on a 400+ model catalog

## At a glance

| Attribute | Subconscious | Fireworks AI |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | DeepSeek V4 Pro, Kimi K3 |
| Speed | 2x faster task completion | 167–174 tok/s on DeepSeek V4 Pro |
| Price | 50–80% lower cost; billed on processed tokens | Fine-tunes served at base price |
| Customization | Marathon post-trained variants | SFT, DPO, RFT; Training API |
| Deployment | Managed API, dedicated, on-prem | Serverless, dedicated GPUs |
| Long context | 5M+ effective context | Full 1M on DeepSeek V4 Pro |

## FAQ

### What is the difference between Subconscious and Fireworks AI?

Fireworks is fast on open-model calls. Subconscious is fast where agents spend their time, on long traces past 200K tokens, and bills only the tokens it processes.

### When should I choose Subconscious over Fireworks AI?

Traces that send 1M tokens but should not bill for all of them; Hour-long coding agents past 200K tokens; Long-horizon agents needing more than 1M tokens of context.

### When should I choose Fireworks AI over Subconscious?

Reinforcement fine-tuning an open model for a narrow task; Teams that want a 400+ model catalog on one API; Latency-sensitive chat and tool calls on a 400+ model catalog.

### Is Subconscious or Fireworks AI cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Fireworks AI: Fine-tunes served at base price. The cheaper choice depends on the model and workload.

### Which has more context, Subconscious or Fireworks AI?

Subconscious: 5M+ effective context. Fireworks AI: Full 1M on DeepSeek V4 Pro.

## Try Subconscious

Subconscious speaks the OpenAI and Anthropic API formats. Base URL: https://api.subconscious.dev/v1. Docs: https://docs.subconscious.dev. Get an API key: https://platform.subconscious.dev/signin. Agent guide: https://www.subconscious.dev/agents.md.

Full profiles: [Subconscious](https://www.subconscious.dev/providers/subconscious.md), [Fireworks AI](https://www.subconscious.dev/providers/fireworks.md).
