We raised $5.1M for long-running agents.
vs

DeepInfra vs Thinking Machines

DeepInfra is the cheapest place to run 150+ open models but offers no managed fine-tuning. Thinking Machines fills exactly that gap with Tinker.

By The Subconscious Team · Updated

DeepInfra vs Thinking Machines: key differences

DeepInfra is the price floor for open inference. Llama 3.1 8B costs $0.02 per million and DeepSeek V4 Flash $0.14 in and $0.28 out, across 150+ models with no minimums or contracts. DeepInfra offers no managed fine-tuning, so training has to happen elsewhere. Thinking Machines is one place it can happen. Tinker's four calls let teams write SFT or RL loops with LoRA adapters on Kimi K2.6, GLM-5.3, Qwen3.5, DeepSeek-V3.1, gpt-oss and Inkling, while the lab runs the distributed GPU work. Pricing is per million tokens across three meters, with GPT-OSS-20B at $0.18 prefill, $0.45 sample and $0.40 train, and cached prefill at 80% off.

Check precision and context on both. DeepInfra serves DeepSeek V4 Pro in FP4, which caps context at 66K, and some reviewers report weaker output unless they pin FP8 variants. Tinker training runs at 32K to 256K context depending on model, while Inkling itself reaches 1M. On serving, DeepInfra is the mature option; Thinking Machines' serverless API is beta and Inkling-only, and checkpoint sampling is scoped to low internal traffic. Bulk tagging, extraction and synthetic data at minimum cost belong on DeepInfra. A team whose base model falls short on a narrow task needs a trainer, and Tinker is built for that.

What DeepInfra and Thinking Machines do

DeepInfra

DeepInfra is the price floor for open-model inference. Developers treat it as the reference point for what a token should cost, with small models like Llama 3.1 8B at $0.02 per million and DeepSeek V4 Flash at $0.14 in and $0.28 out. The catalog covers 150+ open models across text, image and speech behind a fully OpenAI-compatible API. There are no minimums, setup fees or contracts on the shared API.

Example models: DeepSeek V4 Flash, Llama 3.1 8B

Full DeepInfra profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose DeepInfra or Thinking Machines?

DeepInfra

Choose DeepInfra for

  • Lowest per-token cost on high-volume jobs
  • Picking from 150+ open models with no contract
  • Budget backends for consumer chat apps

Thinking Machines

Choose Thinking Machines for

  • Managed-GPU fine-tuning that DeepInfra lacks
  • RL on a custom reward for a narrow task
  • Open Inkling models with 1M context

DeepInfra vs Thinking Machines at a glance

AttributeDeepInfraThinking Machines
Model accessOpen weightsOpen weights
Flagship modelsDeepSeek V4 Flash, Llama 3.1 8BInkling, Inkling-Small
Speed~33 tok/s on DeepSeek V4 Pro (FP4)Unknown
PriceFrom $0.02 per 1MPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationNo managed fine-tuningLoRA SFT and RL via Tinker
DeploymentShared API, no contractsTraining API, beta serverless (Inkling only)
Long context66K on FP4 DeepSeek V4 ProInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between DeepInfra and Thinking Machines?

DeepInfra is the cheapest place to run 150+ open models but offers no managed fine-tuning. Thinking Machines fills exactly that gap with Tinker.

When should I choose DeepInfra over Thinking Machines?

Lowest per-token cost on high-volume jobs; Picking from 150+ open models with no contract; Budget backends for consumer chat apps.

When should I choose Thinking Machines over DeepInfra?

Managed-GPU fine-tuning that DeepInfra lacks; RL on a custom reward for a narrow task; Open Inkling models with 1M context.

Is DeepInfra or Thinking Machines cheaper?

DeepInfra: From $0.02 per 1M. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, DeepInfra or Thinking Machines?

DeepInfra: 66K on FP4 DeepSeek V4 Pro. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.