We raised $5.1M for long-running agents.
vs

Baseten vs Thinking Machines

Baseten is a serving company with the lowest measured time to first token. Thinking Machines is a training company. They are more likely to be used together than against each other.

By The Subconscious Team · Updated

Baseten vs Thinking Machines: key differences

Baseten's Model APIs serve 13 curated open models, including GLM 5.2, DeepSeek V4, Kimi K3 and gpt-oss 120B, over endpoints that speak both OpenAI and Anthropic formats. It posted 0.49 seconds time to first token on the Artificial Analysis board in August 2026, the lowest measured, and dedicated deployments take any model packaged with Truss at about $6.50 an hour on an H100. Baseten's pitch centers on serving rather than training. Thinking Machines is almost entirely training. Tinker lets teams write SFT or RL loops with LoRA adapters on Kimi K2.6, GLM-5.3, Qwen3.5, gpt-oss and Inkling, billed per million prefill, sample and train tokens.

Thinking Machines does serve a little. A beta serverless API covers Inkling at $1.00 in and $4.05 out, and an OpenAI-compatible endpoint samples fine-tuned checkpoints, though the docs scope it to testing and low internal traffic. That leaves production traffic to someone else, which is Baseten's strength: per-minute billing with scale to zero, a 99.99% uptime SLA, KV cache-aware routing for agentic coding, and self-host or HIPAA options. Baseten also runs white-label APIs for model labs. For a team building a custom model, Tinker handles the post-training and Baseten or a similar host handles the serving.

What Baseten and Thinking Machines do

Baseten

Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.

Example models: GLM 5.2, gpt-oss 120B

Full Baseten profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Baseten or Thinking Machines?

Baseten

Choose Baseten for

  • Low-latency production serving of custom models
  • OpenAI and Anthropic-compatible endpoints for coding agents
  • Regulated buyers needing self-host or HIPAA

Thinking Machines

Choose Thinking Machines for

  • Writing custom SFT or RL loops on open bases
  • Training without provisioning GPU clusters
  • Access to the Apache 2.0 Inkling models

Baseten vs Thinking Machines at a glance

AttributeBasetenThinking Machines
Model accessOpen weights, 13 curatedOpen weights
Flagship modelsGLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120BInkling, Inkling-Small
Speed0.49s TTFT, lowest measuredUnknown
PriceH100 about $6.50/hr dedicatedPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationDeploy any model with TrussLoRA SFT and RL via Tinker
DeploymentModel APIs, dedicated, self-hostTraining API, beta serverless (Inkling only)
Long contextVaries by modelInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Baseten and Thinking Machines?

Baseten is a serving company with the lowest measured time to first token. Thinking Machines is a training company. They are more likely to be used together than against each other.

When should I choose Baseten over Thinking Machines?

Low-latency production serving of custom models; OpenAI and Anthropic-compatible endpoints for coding agents; Regulated buyers needing self-host or HIPAA.

When should I choose Thinking Machines over Baseten?

Writing custom SFT or RL loops on open bases; Training without provisioning GPU clusters; Access to the Apache 2.0 Inkling models.

Is Baseten or Thinking Machines cheaper?

Baseten: H100 about $6.50/hr dedicated. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Baseten or Thinking Machines?

Baseten: Varies by model. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.