Thinking Machines vs Infron
Thinking Machines sells Tinker for post-training and the Inkling models. Infron is an inference gateway across 400+ models.
By The Subconscious Team · Updated
Thinking Machines vs Infron: key differences
Thinking Machines' Tinker is an API for LoRA SFT and RL on open-weight models, and its Inkling models run on a beta serverless tier. Infron is a gateway: one OpenAI-compatible API in front of 400+ models from 100+ providers, at provider rates plus a 3% to 5% fee on credit top-ups, with fallbacks, region pinning and a 99.9% uptime SLA on dedicated throughput.
These rarely compete. Thinking Machines is for training; Infron is for calling hosted models. Infron has no fine-tuning, and Tinker is not meant for user-facing traffic.
What Thinking Machines and Infron do
Thinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileInfron
Infron is a US-based AI gateway and inference platform. One OpenAI-compatible API reaches 400+ models from 100+ providers, including DeepSeek, Qwen, Claude, Gemini and GPT through what Infron calls official partner routes, plus media and search models. Teams set provider preferences and fallbacks, see usage and billing in one place, and can bring their own provider keys at no fee. Lawrence Xu is CEO and co-founder Andrew Zheng is CTO.
Example models: DeepSeek, Qwen, Claude, Gemini, GPT
Full Infron profileShould you choose Thinking Machines or Infron?
Thinking Machines
Choose Thinking Machines for
- LoRA SFT and RL via API
- The open Inkling models
- Custom training loops
Infron
Choose Infron for
- Production calls across many vendors
- Automatic failover across providers
- Closed and open models on one key and one bill
Thinking Machines vs Infron at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Closed and open, 400+ models |
| Flagship models | Inkling, Inkling-Small | DeepSeek, Qwen, Claude, Gemini, GPT |
| Speed | Unknown | Unknown |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | Provider rates; 3–5% top-up fee |
| Customization | LoRA SFT and RL via Tinker | Custom deployments |
| Deployment | Training API, beta serverless (Inkling only) | Gateway API, dedicated, BYOK |
| Long context | Inkling up to 1M; Tinker 32K–256K | Varies by model |
Frequently asked questions
What is the difference between Thinking Machines and Infron?
Thinking Machines sells Tinker for post-training and the Inkling models. Infron is an inference gateway across 400+ models.
When should I choose Thinking Machines over Infron?
LoRA SFT and RL via API; The open Inkling models; Custom training loops.
When should I choose Infron over Thinking Machines?
Production calls across many vendors; Automatic failover across providers; Closed and open models on one key and one bill.
Is Thinking Machines or Infron cheaper?
Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Infron: Provider rates; 3–5% top-up fee. The cheaper choice depends on the model and workload.
Which has more context, Thinking Machines or Infron?
Thinking Machines: Inkling up to 1M; Tinker 32K–256K. Infron: Varies by model.
Related comparisons
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Subconscious vs Infron
OpenAI vs Infron
Anthropic vs Infron
Google Vertex AI vs Infron
Amazon Bedrock vs Infron
Together AI vs Infron
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.