DeepSeek vs Thinking Machines
DeepSeek ships MIT-licensed models and a very cheap API. Thinking Machines ships Apache 2.0 Inkling and a way to post-train DeepSeek-V3.1 and other open bases.
By The Subconscious Team · Updated
DeepSeek vs Thinking Machines: key differences
Both labs release open weights, but their businesses differ. DeepSeek's API serves V4.1 Flash at $0.30 in and $1.20 out at peak and V4 Pro at $1.32 in and $3.96 out, both with 1M context and 384K max output, with every off-peak hour at half price and cache hits costing a few cents per million. Weights ship under MIT. Thinking Machines sells training first. Tinker lets teams write SFT or RL loops with LoRA adapters on open bases, including DeepSeek-V3.1, Kimi K2.6, GLM-5.3 and Qwen3.5. Its own Inkling models are Apache 2.0, with Inkling at $1.00 in and $4.05 out on a beta serverless API and up to 1M context.
For raw inference, DeepSeek is cheaper and more established, though hosted API data is stored in China, which stops many enterprises, and frequent repricing means cost models need rechecking. Thinking Machines is a US lab, and Inkling adds native image and audio input where V4.1 Flash has built-in image understanding. Its serving is limited: only the two Inkling models, in beta, and checkpoint sampling scoped to testing. DeepSeek's open weights can be fine-tuned anywhere; Tinker is one managed way to do it without running GPUs. Cost-first agents scheduled off-peak favor DeepSeek. Teams building a specialized model favor Tinker.
What DeepSeek and Thinking Machines do
DeepSeek
DeepSeek is the Chinese lab whose open-weight models reset price expectations for the whole market. Its API now serves two models, both with 1M context and 384K max output. V4.1 Flash shipped September 10, 2026 with built-in image understanding at $0.30 in and $1.20 out at peak. V4 Pro, generally available since August 13, costs $1.32 in and $3.96 out at peak. Cache hits cost a few cents per million or less, and the weights ship on Hugging Face under an MIT license.
Example models: DeepSeek V4.1 Flash, DeepSeek V4 Pro
Full DeepSeek profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose DeepSeek or Thinking Machines?
DeepSeek
Choose DeepSeek for
- Cheapest first-party API for strong open models
- Batch work scheduled into off-peak hours
- Long outputs up to 384K tokens
Thinking Machines
Choose Thinking Machines for
- Managed post-training of DeepSeek-V3.1 and others
- A US-based lab for teams avoiding China-hosted data
- Inkling with native audio input
DeepSeek vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights (MIT) | Open weights |
| Flagship models | DeepSeek V4.1 Flash, V4 Pro | Inkling, Inkling-Small |
| Speed | ~35 tok/s on V4 Pro | Unknown |
| Price | Off-peak hours at half price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Open weights to fine-tune | LoRA SFT and RL via Tinker |
| Deployment | First-party API, Hugging Face weights | Training API, beta serverless (Inkling only) |
| Long context | 1M, 384K max output | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between DeepSeek and Thinking Machines?
DeepSeek ships MIT-licensed models and a very cheap API. Thinking Machines ships Apache 2.0 Inkling and a way to post-train DeepSeek-V3.1 and other open bases.
When should I choose DeepSeek over Thinking Machines?
Cheapest first-party API for strong open models; Batch work scheduled into off-peak hours; Long outputs up to 384K tokens.
When should I choose Thinking Machines over DeepSeek?
Managed post-training of DeepSeek-V3.1 and others; A US-based lab for teams avoiding China-hosted data; Inkling with native audio input.
Is DeepSeek or Thinking Machines cheaper?
DeepSeek: Off-peak hours at half price. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, DeepSeek or Thinking Machines?
DeepSeek: 1M, 384K max output. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs DeepSeek
OpenAI vs DeepSeek
Anthropic vs DeepSeek
Google Vertex AI vs DeepSeek
Amazon Bedrock vs DeepSeek
Together AI vs DeepSeek
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.