Baseten vs Thinking Machines
Baseten is a serving company with the lowest measured time to first token. Thinking Machines is a training company. They are more likely to be used together than against each other.
By The Subconscious Team · Updated
Baseten vs Thinking Machines: key differences
Baseten's Model APIs serve 13 curated open models, including GLM 5.2, DeepSeek V4, Kimi K3 and gpt-oss 120B, over endpoints that speak both OpenAI and Anthropic formats. It posted 0.49 seconds time to first token on the Artificial Analysis board in August 2026, the lowest measured, and dedicated deployments take any model packaged with Truss at about $6.50 an hour on an H100. Baseten's pitch centers on serving rather than training. Thinking Machines is almost entirely training. Tinker lets teams write SFT or RL loops with LoRA adapters on Kimi K2.6, GLM-5.3, Qwen3.5, gpt-oss and Inkling, billed per million prefill, sample and train tokens.
Thinking Machines does serve a little. A beta serverless API covers Inkling at $1.00 in and $4.05 out, and an OpenAI-compatible endpoint samples fine-tuned checkpoints, though the docs scope it to testing and low internal traffic. That leaves production traffic to someone else, which is Baseten's strength: per-minute billing with scale to zero, a 99.99% uptime SLA, KV cache-aware routing for agentic coding, and self-host or HIPAA options. Baseten also runs white-label APIs for model labs. For a team building a custom model, Tinker handles the post-training and Baseten or a similar host handles the serving.
What Baseten and Thinking Machines do
Baseten
Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.
Example models: GLM 5.2, gpt-oss 120B
Full Baseten profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Baseten or Thinking Machines?
Baseten
Choose Baseten for
- Low-latency production serving of custom models
- OpenAI and Anthropic-compatible endpoints for coding agents
- Regulated buyers needing self-host or HIPAA
Thinking Machines
Choose Thinking Machines for
- Writing custom SFT or RL loops on open bases
- Training without provisioning GPU clusters
- Access to the Apache 2.0 Inkling models
Baseten vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights, 13 curated | Open weights |
| Flagship models | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | Inkling, Inkling-Small |
| Speed | 0.49s TTFT, lowest measured | Unknown |
| Price | H100 about $6.50/hr dedicated | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Deploy any model with Truss | LoRA SFT and RL via Tinker |
| Deployment | Model APIs, dedicated, self-host | Training API, beta serverless (Inkling only) |
| Long context | Varies by model | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Baseten and Thinking Machines?
Baseten is a serving company with the lowest measured time to first token. Thinking Machines is a training company. They are more likely to be used together than against each other.
When should I choose Baseten over Thinking Machines?
Low-latency production serving of custom models; OpenAI and Anthropic-compatible endpoints for coding agents; Regulated buyers needing self-host or HIPAA.
When should I choose Thinking Machines over Baseten?
Writing custom SFT or RL loops on open bases; Training without provisioning GPU clusters; Access to the Apache 2.0 Inkling models.
Is Baseten or Thinking Machines cheaper?
Baseten: H100 about $6.50/hr dedicated. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Baseten or Thinking Machines?
Baseten: Varies by model. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Baseten
OpenAI vs Baseten
Anthropic vs Baseten
Google Vertex AI vs Baseten
Amazon Bedrock vs Baseten
Together AI vs Baseten
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.