# Baseten vs Thinking Machines

> Baseten is a serving company with the lowest measured time to first token. Thinking Machines is a training company. They are more likely to be used together than against each other.

Canonical: https://www.subconscious.dev/compare/baseten-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Baseten's Model APIs serve 13 curated open models, including GLM 5.2, DeepSeek V4, Kimi K3 and gpt-oss 120B, over endpoints that speak both OpenAI and Anthropic formats. It posted 0.49 seconds time to first token on the Artificial Analysis board in August 2026, the lowest measured, and dedicated deployments take any model packaged with Truss at about $6.50 an hour on an H100. Baseten's pitch centers on serving rather than training. Thinking Machines is almost entirely training. Tinker lets teams write SFT or RL loops with LoRA adapters on Kimi K2.6, GLM-5.3, Qwen3.5, gpt-oss and Inkling, billed per million prefill, sample and train tokens.

Thinking Machines does serve a little. A beta serverless API covers Inkling at $1.00 in and $4.05 out, and an OpenAI-compatible endpoint samples fine-tuned checkpoints, though the docs scope it to testing and low internal traffic. That leaves production traffic to someone else, which is Baseten's strength: per-minute billing with scale to zero, a 99.99% uptime SLA, KV cache-aware routing for agentic coding, and self-host or HIPAA options. Baseten also runs white-label APIs for model labs. For a team building a custom model, Tinker handles the post-training and Baseten or a similar host handles the serving.

## What each one does

### Baseten

Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Baseten for

- Low-latency production serving of custom models
- OpenAI and Anthropic-compatible endpoints for coding agents
- Regulated buyers needing self-host or HIPAA

### Choose Thinking Machines for

- Writing custom SFT or RL loops on open bases
- Training without provisioning GPU clusters
- Access to the Apache 2.0 Inkling models

## At a glance

| Attribute | Baseten | Thinking Machines |
|---|---|---|
| Model access | Open weights, 13 curated | Open weights |
| Flagship models | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | Inkling, Inkling-Small |
| Speed | 0.49s TTFT, lowest measured | - |
| Price | H100 about $6.50/hr dedicated | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Deploy any model with Truss | LoRA SFT and RL via Tinker |
| Deployment | Model APIs, dedicated, self-host | Training API, beta serverless (Inkling only) |
| Long context | Varies by model | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Baseten and Thinking Machines?

Baseten is a serving company with the lowest measured time to first token. Thinking Machines is a training company. They are more likely to be used together than against each other.

### When should I choose Baseten over Thinking Machines?

Low-latency production serving of custom models; OpenAI and Anthropic-compatible endpoints for coding agents; Regulated buyers needing self-host or HIPAA.

### When should I choose Thinking Machines over Baseten?

Writing custom SFT or RL loops on open bases; Training without provisioning GPU clusters; Access to the Apache 2.0 Inkling models.

### Is Baseten or Thinking Machines cheaper?

Baseten: H100 about $6.50/hr dedicated. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Baseten or Thinking Machines?

Baseten: Varies by model. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Baseten](https://www.subconscious.dev/compare/subconscious-vs-baseten.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Baseten](https://www.subconscious.dev/providers/baseten.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
