Thinking Machines vs Sail Research
Thinking Machines trains open models through Tinker. Sail Research serves them cheaply by letting customers wait, with completion windows at 30 to 80% off.
By The Subconscious Team · Updated
Thinking Machines vs Sail Research: key differences
Both reach for large open models like Kimi K2.6 and GPT-OSS, but for different jobs. Sail Research serves inference on a throughput-first stack. Customers pick a window: priority targets about a minute per turn for 30 to 50% off, standard about five minutes for 45 to 65% off, and flex runs off-peak for 60 to 80% off. It serves customer LoRA fine-tunes and gives agents persistent Sailboxes. Thinking Machines sells the training side. Tinker runs LoRA SFT and RL loops the customer writes, billed on prefill, sample and train meters, with cached prefill 80% off.
A team could train an adapter on Tinker and run it on Sail for hours-long background agents, assuming a supported base model. Sail claims 3x to 10x savings over comparable hosts. It is explicitly unsuited to live chat or voice, and it serves only open models. Thinking Machines' serving is thinner still: a beta serverless API for Inkling at $1.00 in and $4.05 out, plus a checkpoint endpoint for testing. Inkling itself stands out, with 1M context and native image and audio input under Apache 2.0. Sail's catalog adds GLM-5, Qwen 3.6 and Gemma 4, and its context varies by model.
What Thinking Machines and Sail Research do
Thinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileSail Research
Sail Research sells throughput over latency. Founders Neil Movva and Samir Menon built a serving stack that packs as much work as possible into every GPU, and customers state how long they can wait through completion windows. The priority window targets about a one-minute turn for roughly 30 to 50% off the immediate asap price. The default standard window targets about five minutes for 45 to 65% off. The flex window runs off-peak for 60 to 80% off.
Example models: Kimi K2.6, GLM-5
Full Sail Research profileShould you choose Thinking Machines or Sail Research?
Thinking Machines
Choose Thinking Machines for
- Custom RL and SFT post-training
- Training LoRA adapters on large MoE bases
- Evaluating Inkling's 1M-context multimodal input
Sail Research
Choose Sail Research for
- Cheap inference for background agents
- Evals and batch jobs that can wait minutes
- Serving LoRA fine-tunes at discounted rates
Thinking Machines vs Sail Research at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | Inkling, Inkling-Small | Kimi K2.6, GLM-5, GPT-OSS 120B |
| Speed | Unknown | Minutes per turn by design |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | 30–80% off by completion window |
| Customization | LoRA SFT and RL via Tinker | Customer LoRA fine-tunes |
| Deployment | Training API, beta serverless (Inkling only) | API plus Sailboxes |
| Long context | Inkling up to 1M; Tinker 32K–256K | Varies by model |
Frequently asked questions
What is the difference between Thinking Machines and Sail Research?
Thinking Machines trains open models through Tinker. Sail Research serves them cheaply by letting customers wait, with completion windows at 30 to 80% off.
When should I choose Thinking Machines over Sail Research?
Custom RL and SFT post-training; Training LoRA adapters on large MoE bases; Evaluating Inkling's 1M-context multimodal input.
When should I choose Sail Research over Thinking Machines?
Cheap inference for background agents; Evals and batch jobs that can wait minutes; Serving LoRA fine-tunes at discounted rates.
Is Thinking Machines or Sail Research cheaper?
Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Sail Research: 30–80% off by completion window. The cheaper choice depends on the model and workload.
Which has more context, Thinking Machines or Sail Research?
Thinking Machines: Inkling up to 1M; Tinker 32K–256K. Sail Research: Varies by model.
Related comparisons
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Subconscious vs Sail Research
OpenAI vs Sail Research
Anthropic vs Sail Research
Google Vertex AI vs Sail Research
Amazon Bedrock vs Sail Research
Together AI vs Sail Research
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.