We raised $5.1M for long-running agents.
vs

Inference.net vs Thinking Machines

Both help teams build custom models. Inference.net turns production traces into a distilled model it hosts. Thinking Machines hands researchers a low-level training API and leaves serving to others.

By The Subconscious Team · Updated

Inference.net vs Thinking Machines: key differences

The difference is how much of the loop each owns. Inference.net runs a managed pipeline: its gateway routes and logs traffic across open, closed and custom models, turns that traffic into eval and training sets, fine-tunes a task-specific model and deploys it on a dedicated GPU with a 99.99% uptime target. It also sells cheap batch on spare GPU capacity, up to 1M requests per file. Thinking Machines gives more control and less hand-holding. Tinker exposes forward_backward, optim_step, sample and save_state, so teams write their own SFT or RL code while the lab runs the distributed LoRA training.

Model scope differs too. Tinker trains large bases like Kimi K2.6, GLM-5.3, DeepSeek-V3.1 and the 975B Inkling, which suits research groups pushing capability. Inference.net aims at replacing a narrow GPT-class workload with a smaller distilled model to cut cost and latency. Serving is Inference.net's clear win. Tinker's checkpoint endpoint is scoped to testing and low internal traffic, and its beta serverless API covers only Inkling at $1.00 in and $4.05 out. Inference.net publishes few independent benchmarks, so buyers lean on its numbers. Thinking Machines publishes per-meter prices, such as GPT-OSS-20B at $0.18 prefill, $0.45 sample and $0.40 train.

What Inference.net and Thinking Machines do

Inference.net

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

Example models: open catalog models plus customer fine-tunes served on dedicated GPUs

Full Inference.net profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Inference.net or Thinking Machines?

Inference.net

Choose Inference.net for

  • Distilling production traces into a hosted model
  • Cheap bulk jobs on spare GPU capacity
  • One gateway for open, closed and custom models

Thinking Machines

Choose Thinking Machines for

  • Hands-on RL research on open weights
  • Post-training large MoE bases
  • Published per-token training prices

Inference.net vs Thinking Machines at a glance

AttributeInference.netThinking Machines
Model accessOpen, closed and customOpen weights
Flagship modelsCustomer fine-tunesInkling, Inkling-Small
SpeedBatch windows of 24h to 7 daysUnknown
PriceDiscounted spare GPU capacityPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationDistill traces into custom modelsLoRA SFT and RL via Tinker
DeploymentBatch API, gateway, dedicated GPUsTraining API, beta serverless (Inkling only)
Long contextVaries by modelInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Inference.net and Thinking Machines?

Both help teams build custom models. Inference.net turns production traces into a distilled model it hosts. Thinking Machines hands researchers a low-level training API and leaves serving to others.

When should I choose Inference.net over Thinking Machines?

Distilling production traces into a hosted model; Cheap bulk jobs on spare GPU capacity; One gateway for open, closed and custom models.

When should I choose Thinking Machines over Inference.net?

Hands-on RL research on open weights; Post-training large MoE bases; Published per-token training prices.

Is Inference.net or Thinking Machines cheaper?

Inference.net: Discounted spare GPU capacity. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Inference.net or Thinking Machines?

Inference.net: Varies by model. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.