# DeepSeek vs Thinking Machines

> DeepSeek ships MIT-licensed models and a very cheap API. Thinking Machines ships Apache 2.0 Inkling and a way to post-train DeepSeek-V3.1 and other open bases.

Canonical: https://www.subconscious.dev/compare/deepseek-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Both labs release open weights, but their businesses differ. DeepSeek's API serves V4.1 Flash at $0.30 in and $1.20 out at peak and V4 Pro at $1.32 in and $3.96 out, both with 1M context and 384K max output, with every off-peak hour at half price and cache hits costing a few cents per million. Weights ship under MIT. Thinking Machines sells training first. Tinker lets teams write SFT or RL loops with LoRA adapters on open bases, including DeepSeek-V3.1, Kimi K2.6, GLM-5.3 and Qwen3.5. Its own Inkling models are Apache 2.0, with Inkling at $1.00 in and $4.05 out on a beta serverless API and up to 1M context.

For raw inference, DeepSeek is cheaper and more established, though hosted API data is stored in China, which stops many enterprises, and frequent repricing means cost models need rechecking. Thinking Machines is a US lab, and Inkling adds native image and audio input where V4.1 Flash has built-in image understanding. Its serving is limited: only the two Inkling models, in beta, and checkpoint sampling scoped to testing. DeepSeek's open weights can be fine-tuned anywhere; Tinker is one managed way to do it without running GPUs. Cost-first agents scheduled off-peak favor DeepSeek. Teams building a specialized model favor Tinker.

## What each one does

### DeepSeek

DeepSeek is the Chinese lab whose open-weight models reset price expectations for the whole market. Its API now serves two models, both with 1M context and 384K max output. V4.1 Flash shipped September 10, 2026 with built-in image understanding at $0.30 in and $1.20 out at peak. V4 Pro, generally available since August 13, costs $1.32 in and $3.96 out at peak. Cache hits cost a few cents per million or less, and the weights ship on Hugging Face under an MIT license.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose DeepSeek for

- Cheapest first-party API for strong open models
- Batch work scheduled into off-peak hours
- Long outputs up to 384K tokens

### Choose Thinking Machines for

- Managed post-training of DeepSeek-V3.1 and others
- A US-based lab for teams avoiding China-hosted data
- Inkling with native audio input

## At a glance

| Attribute | DeepSeek | Thinking Machines |
|---|---|---|
| Model access | Open weights (MIT) | Open weights |
| Flagship models | DeepSeek V4.1 Flash, V4 Pro | Inkling, Inkling-Small |
| Speed | ~35 tok/s on V4 Pro | - |
| Price | Off-peak hours at half price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Open weights to fine-tune | LoRA SFT and RL via Tinker |
| Deployment | First-party API, Hugging Face weights | Training API, beta serverless (Inkling only) |
| Long context | 1M, 384K max output | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between DeepSeek and Thinking Machines?

DeepSeek ships MIT-licensed models and a very cheap API. Thinking Machines ships Apache 2.0 Inkling and a way to post-train DeepSeek-V3.1 and other open bases.

### When should I choose DeepSeek over Thinking Machines?

Cheapest first-party API for strong open models; Batch work scheduled into off-peak hours; Long outputs up to 384K tokens.

### When should I choose Thinking Machines over DeepSeek?

Managed post-training of DeepSeek-V3.1 and others; A US-based lab for teams avoiding China-hosted data; Inkling with native audio input.

### Is DeepSeek or Thinking Machines cheaper?

DeepSeek: Off-peak hours at half price. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, DeepSeek or Thinking Machines?

DeepSeek: 1M, 384K max output. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs DeepSeek](https://www.subconscious.dev/compare/subconscious-vs-deepseek.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [DeepSeek](https://www.subconscious.dev/providers/deepseek.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
