# DeepInfra vs Thinking Machines

> DeepInfra is the cheapest place to run 150+ open models but offers no managed fine-tuning. Thinking Machines fills exactly that gap with Tinker.

Canonical: https://www.subconscious.dev/compare/deepinfra-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

DeepInfra is the price floor for open inference. Llama 3.1 8B costs $0.02 per million and DeepSeek V4 Flash $0.14 in and $0.28 out, across 150+ models with no minimums or contracts. DeepInfra offers no managed fine-tuning, so training has to happen elsewhere. Thinking Machines is one place it can happen. Tinker's four calls let teams write SFT or RL loops with LoRA adapters on Kimi K2.6, GLM-5.3, Qwen3.5, DeepSeek-V3.1, gpt-oss and Inkling, while the lab runs the distributed GPU work. Pricing is per million tokens across three meters, with GPT-OSS-20B at $0.18 prefill, $0.45 sample and $0.40 train, and cached prefill at 80% off.

Check precision and context on both. DeepInfra serves DeepSeek V4 Pro in FP4, which caps context at 66K, and some reviewers report weaker output unless they pin FP8 variants. Tinker training runs at 32K to 256K context depending on model, while Inkling itself reaches 1M. On serving, DeepInfra is the mature option; Thinking Machines' serverless API is beta and Inkling-only, and checkpoint sampling is scoped to low internal traffic. Bulk tagging, extraction and synthetic data at minimum cost belong on DeepInfra. A team whose base model falls short on a narrow task needs a trainer, and Tinker is built for that.

## What each one does

### DeepInfra

DeepInfra is the price floor for open-model inference. Developers treat it as the reference point for what a token should cost, with small models like Llama 3.1 8B at $0.02 per million and DeepSeek V4 Flash at $0.14 in and $0.28 out. The catalog covers 150+ open models across text, image and speech behind a fully OpenAI-compatible API. There are no minimums, setup fees or contracts on the shared API.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose DeepInfra for

- Lowest per-token cost on high-volume jobs
- Picking from 150+ open models with no contract
- Budget backends for consumer chat apps

### Choose Thinking Machines for

- Managed-GPU fine-tuning that DeepInfra lacks
- RL on a custom reward for a narrow task
- Open Inkling models with 1M context

## At a glance

| Attribute | DeepInfra | Thinking Machines |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | DeepSeek V4 Flash, Llama 3.1 8B | Inkling, Inkling-Small |
| Speed | ~33 tok/s on DeepSeek V4 Pro (FP4) | - |
| Price | From $0.02 per 1M | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | No managed fine-tuning | LoRA SFT and RL via Tinker |
| Deployment | Shared API, no contracts | Training API, beta serverless (Inkling only) |
| Long context | 66K on FP4 DeepSeek V4 Pro | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between DeepInfra and Thinking Machines?

DeepInfra is the cheapest place to run 150+ open models but offers no managed fine-tuning. Thinking Machines fills exactly that gap with Tinker.

### When should I choose DeepInfra over Thinking Machines?

Lowest per-token cost on high-volume jobs; Picking from 150+ open models with no contract; Budget backends for consumer chat apps.

### When should I choose Thinking Machines over DeepInfra?

Managed-GPU fine-tuning that DeepInfra lacks; RL on a custom reward for a narrow task; Open Inkling models with 1M context.

### Is DeepInfra or Thinking Machines cheaper?

DeepInfra: From $0.02 per 1M. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, DeepInfra or Thinking Machines?

DeepInfra: 66K on FP4 DeepSeek V4 Pro. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs DeepInfra](https://www.subconscious.dev/compare/subconscious-vs-deepinfra.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [DeepInfra](https://www.subconscious.dev/providers/deepinfra.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
