Inference.net vs Thinking Machines
Both help teams build custom models. Inference.net turns production traces into a distilled model it hosts. Thinking Machines hands researchers a low-level training API and leaves serving to others.
By The Subconscious Team · Updated
Inference.net vs Thinking Machines: key differences
The difference is how much of the loop each owns. Inference.net runs a managed pipeline: its gateway routes and logs traffic across open, closed and custom models, turns that traffic into eval and training sets, fine-tunes a task-specific model and deploys it on a dedicated GPU with a 99.99% uptime target. It also sells cheap batch on spare GPU capacity, up to 1M requests per file. Thinking Machines gives more control and less hand-holding. Tinker exposes forward_backward, optim_step, sample and save_state, so teams write their own SFT or RL code while the lab runs the distributed LoRA training.
Model scope differs too. Tinker trains large bases like Kimi K2.6, GLM-5.3, DeepSeek-V3.1 and the 975B Inkling, which suits research groups pushing capability. Inference.net aims at replacing a narrow GPT-class workload with a smaller distilled model to cut cost and latency. Serving is Inference.net's clear win. Tinker's checkpoint endpoint is scoped to testing and low internal traffic, and its beta serverless API covers only Inkling at $1.00 in and $4.05 out. Inference.net publishes few independent benchmarks, so buyers lean on its numbers. Thinking Machines publishes per-meter prices, such as GPT-OSS-20B at $0.18 prefill, $0.45 sample and $0.40 train.
What Inference.net and Thinking Machines do
Inference.net
Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.
Example models: open catalog models plus customer fine-tunes served on dedicated GPUs
Full Inference.net profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Inference.net or Thinking Machines?
Inference.net
Choose Inference.net for
- Distilling production traces into a hosted model
- Cheap bulk jobs on spare GPU capacity
- One gateway for open, closed and custom models
Thinking Machines
Choose Thinking Machines for
- Hands-on RL research on open weights
- Post-training large MoE bases
- Published per-token training prices
Inference.net vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open, closed and custom | Open weights |
| Flagship models | Customer fine-tunes | Inkling, Inkling-Small |
| Speed | Batch windows of 24h to 7 days | Unknown |
| Price | Discounted spare GPU capacity | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Distill traces into custom models | LoRA SFT and RL via Tinker |
| Deployment | Batch API, gateway, dedicated GPUs | Training API, beta serverless (Inkling only) |
| Long context | Varies by model | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Inference.net and Thinking Machines?
Both help teams build custom models. Inference.net turns production traces into a distilled model it hosts. Thinking Machines hands researchers a low-level training API and leaves serving to others.
When should I choose Inference.net over Thinking Machines?
Distilling production traces into a hosted model; Cheap bulk jobs on spare GPU capacity; One gateway for open, closed and custom models.
When should I choose Thinking Machines over Inference.net?
Hands-on RL research on open weights; Post-training large MoE bases; Published per-token training prices.
Is Inference.net or Thinking Machines cheaper?
Inference.net: Discounted spare GPU capacity. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Inference.net or Thinking Machines?
Inference.net: Varies by model. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Inference.net
OpenAI vs Inference.net
Anthropic vs Inference.net
Google Vertex AI vs Inference.net
Amazon Bedrock vs Inference.net
Together AI vs Inference.net
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.