We raised $5.1M for long-running agents.
vs

DeepSeek vs Thinking Machines

DeepSeek ships MIT-licensed models and a very cheap API. Thinking Machines ships Apache 2.0 Inkling and a way to post-train DeepSeek-V3.1 and other open bases.

By The Subconscious Team · Updated

DeepSeek vs Thinking Machines: key differences

Both labs release open weights, but their businesses differ. DeepSeek's API serves V4.1 Flash at $0.30 in and $1.20 out at peak and V4 Pro at $1.32 in and $3.96 out, both with 1M context and 384K max output, with every off-peak hour at half price and cache hits costing a few cents per million. Weights ship under MIT. Thinking Machines sells training first. Tinker lets teams write SFT or RL loops with LoRA adapters on open bases, including DeepSeek-V3.1, Kimi K2.6, GLM-5.3 and Qwen3.5. Its own Inkling models are Apache 2.0, with Inkling at $1.00 in and $4.05 out on a beta serverless API and up to 1M context.

For raw inference, DeepSeek is cheaper and more established, though hosted API data is stored in China, which stops many enterprises, and frequent repricing means cost models need rechecking. Thinking Machines is a US lab, and Inkling adds native image and audio input where V4.1 Flash has built-in image understanding. Its serving is limited: only the two Inkling models, in beta, and checkpoint sampling scoped to testing. DeepSeek's open weights can be fine-tuned anywhere; Tinker is one managed way to do it without running GPUs. Cost-first agents scheduled off-peak favor DeepSeek. Teams building a specialized model favor Tinker.

What DeepSeek and Thinking Machines do

DeepSeek

DeepSeek is the Chinese lab whose open-weight models reset price expectations for the whole market. Its API now serves two models, both with 1M context and 384K max output. V4.1 Flash shipped September 10, 2026 with built-in image understanding at $0.30 in and $1.20 out at peak. V4 Pro, generally available since August 13, costs $1.32 in and $3.96 out at peak. Cache hits cost a few cents per million or less, and the weights ship on Hugging Face under an MIT license.

Example models: DeepSeek V4.1 Flash, DeepSeek V4 Pro

Full DeepSeek profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose DeepSeek or Thinking Machines?

DeepSeek

Choose DeepSeek for

  • Cheapest first-party API for strong open models
  • Batch work scheduled into off-peak hours
  • Long outputs up to 384K tokens

Thinking Machines

Choose Thinking Machines for

  • Managed post-training of DeepSeek-V3.1 and others
  • A US-based lab for teams avoiding China-hosted data
  • Inkling with native audio input

DeepSeek vs Thinking Machines at a glance

AttributeDeepSeekThinking Machines
Model accessOpen weights (MIT)Open weights
Flagship modelsDeepSeek V4.1 Flash, V4 ProInkling, Inkling-Small
Speed~35 tok/s on V4 ProUnknown
PriceOff-peak hours at half pricePer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationOpen weights to fine-tuneLoRA SFT and RL via Tinker
DeploymentFirst-party API, Hugging Face weightsTraining API, beta serverless (Inkling only)
Long context1M, 384K max outputInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between DeepSeek and Thinking Machines?

DeepSeek ships MIT-licensed models and a very cheap API. Thinking Machines ships Apache 2.0 Inkling and a way to post-train DeepSeek-V3.1 and other open bases.

When should I choose DeepSeek over Thinking Machines?

Cheapest first-party API for strong open models; Batch work scheduled into off-peak hours; Long outputs up to 384K tokens.

When should I choose Thinking Machines over DeepSeek?

Managed post-training of DeepSeek-V3.1 and others; A US-based lab for teams avoiding China-hosted data; Inkling with native audio input.

Is DeepSeek or Thinking Machines cheaper?

DeepSeek: Off-peak hours at half price. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, DeepSeek or Thinking Machines?

DeepSeek: 1M, 384K max output. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.