# Fireworks AI vs Thinking Machines

> Fireworks and Thinking Machines both sell researcher-grade RL training on open models. Fireworks also runs some of the fastest GPU serving; Thinking Machines does not yet serve at scale.

Canonical: https://www.subconscious.dev/compare/fireworks-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Training is where these two collide. Fireworks offers SFT, DPO and reinforcement fine-tuning in LoRA or full-parameter form, and its Training API, GA since August 31, 2026, lets researchers run their own RL loop on Fireworks-managed trainers with matched numerics between training and inference. Thinking Machines' Tinker has been doing the low-level version since October 2025, with four calls that let teams write any SFT or RL loop, but it trains LoRA adapters only. Tinker bills per million prefill, sample and train tokens. Fireworks serves a fine-tuned model at the same per-token price as its base, which removes the usual markup once a model is ready.

Serving is the widest gap. Fireworks has posted 167 to 174 tokens per second on DeepSeek V4 Pro in third-party measurements, serves the full 1M context there, and carries 400+ models with SOC 2, HIPAA and ISO certifications. Thinking Machines' serverless API is beta and covers only Inkling and Inkling-Small, and checkpoint sampling is scoped to testing. Thinking Machines' strengths are its own Apache 2.0 Inkling models, with native image and audio input and 1M context, and support for bases like Kimi K2.6 and GLM-5.3. Fireworks fits teams that train and then ship; Tinker fits research groups focused on the training loop itself.

## What each one does

### Fireworks AI

Fireworks AI was founded in 2022 by former Meta PyTorch engineers led by CEO Lin Qiao, and it sells speed on open models. Its custom serving stack has posted 167 to 174 tokens per second on DeepSeek V4 Pro in third-party measurements, several times what most GPU peers hit on the same model. The catalog holds 400+ models across text, vision, audio and embeddings, served through an OpenAI-compatible API. In July 2026 it raised a $1.505B Series D at a $17.5B valuation, with a reported $1B+ run rate and 40T+ tokens a day.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Fireworks AI for

- RL fine-tunes that go straight to fast production serving
- Full-parameter training, not only LoRA
- Compliance needs like HIPAA and SOC 2

### Choose Thinking Machines for

- Low-level loop control with a small four-call API
- Post-training Inkling with image and audio input
- Research runs sampled mid-training

## At a glance

| Attribute | Fireworks AI | Thinking Machines |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | DeepSeek V4 Pro, Kimi K3 | Inkling, Inkling-Small |
| Speed | 167–174 tok/s on DeepSeek V4 Pro | - |
| Price | Fine-tunes served at base price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | SFT, DPO, RFT; Training API | LoRA SFT and RL via Tinker |
| Deployment | Serverless, dedicated GPUs | Training API, beta serverless (Inkling only) |
| Long context | Full 1M on DeepSeek V4 Pro | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Fireworks AI and Thinking Machines?

Fireworks and Thinking Machines both sell researcher-grade RL training on open models. Fireworks also runs some of the fastest GPU serving; Thinking Machines does not yet serve at scale.

### When should I choose Fireworks AI over Thinking Machines?

RL fine-tunes that go straight to fast production serving; Full-parameter training, not only LoRA; Compliance needs like HIPAA and SOC 2.

### When should I choose Thinking Machines over Fireworks AI?

Low-level loop control with a small four-call API; Post-training Inkling with image and audio input; Research runs sampled mid-training.

### Is Fireworks AI or Thinking Machines cheaper?

Fireworks AI: Fine-tunes served at base price. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Fireworks AI or Thinking Machines?

Fireworks AI: Full 1M on DeepSeek V4 Pro. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Fireworks AI](https://www.subconscious.dev/compare/subconscious-vs-fireworks.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Fireworks AI](https://www.subconscious.dev/providers/fireworks.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
