We raised $5.1M for long-running agents.
vs

Thinking Machines vs Sail Research

Thinking Machines trains open models through Tinker. Sail Research serves them cheaply by letting customers wait, with completion windows at 30 to 80% off.

By The Subconscious Team · Updated

Thinking Machines vs Sail Research: key differences

Both reach for large open models like Kimi K2.6 and GPT-OSS, but for different jobs. Sail Research serves inference on a throughput-first stack. Customers pick a window: priority targets about a minute per turn for 30 to 50% off, standard about five minutes for 45 to 65% off, and flex runs off-peak for 60 to 80% off. It serves customer LoRA fine-tunes and gives agents persistent Sailboxes. Thinking Machines sells the training side. Tinker runs LoRA SFT and RL loops the customer writes, billed on prefill, sample and train meters, with cached prefill 80% off.

A team could train an adapter on Tinker and run it on Sail for hours-long background agents, assuming a supported base model. Sail claims 3x to 10x savings over comparable hosts. It is explicitly unsuited to live chat or voice, and it serves only open models. Thinking Machines' serving is thinner still: a beta serverless API for Inkling at $1.00 in and $4.05 out, plus a checkpoint endpoint for testing. Inkling itself stands out, with 1M context and native image and audio input under Apache 2.0. Sail's catalog adds GLM-5, Qwen 3.6 and Gemma 4, and its context varies by model.

What Thinking Machines and Sail Research do

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Sail Research

Sail Research sells throughput over latency. Founders Neil Movva and Samir Menon built a serving stack that packs as much work as possible into every GPU, and customers state how long they can wait through completion windows. The priority window targets about a one-minute turn for roughly 30 to 50% off the immediate asap price. The default standard window targets about five minutes for 45 to 65% off. The flex window runs off-peak for 60 to 80% off.

Example models: Kimi K2.6, GLM-5

Full Sail Research profile

Should you choose Thinking Machines or Sail Research?

Thinking Machines

Choose Thinking Machines for

  • Custom RL and SFT post-training
  • Training LoRA adapters on large MoE bases
  • Evaluating Inkling's 1M-context multimodal input

Sail Research

Choose Sail Research for

  • Cheap inference for background agents
  • Evals and batch jobs that can wait minutes
  • Serving LoRA fine-tunes at discounted rates

Thinking Machines vs Sail Research at a glance

AttributeThinking MachinesSail Research
Model accessOpen weightsOpen weights
Flagship modelsInkling, Inkling-SmallKimi K2.6, GLM-5, GPT-OSS 120B
SpeedUnknownMinutes per turn by design
PricePer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out30–80% off by completion window
CustomizationLoRA SFT and RL via TinkerCustomer LoRA fine-tunes
DeploymentTraining API, beta serverless (Inkling only)API plus Sailboxes
Long contextInkling up to 1M; Tinker 32K–256KVaries by model

Frequently asked questions

What is the difference between Thinking Machines and Sail Research?

Thinking Machines trains open models through Tinker. Sail Research serves them cheaply by letting customers wait, with completion windows at 30 to 80% off.

When should I choose Thinking Machines over Sail Research?

Custom RL and SFT post-training; Training LoRA adapters on large MoE bases; Evaluating Inkling's 1M-context multimodal input.

When should I choose Sail Research over Thinking Machines?

Cheap inference for background agents; Evals and batch jobs that can wait minutes; Serving LoRA fine-tunes at discounted rates.

Is Thinking Machines or Sail Research cheaper?

Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Sail Research: 30–80% off by completion window. The cheaper choice depends on the model and workload.

Which has more context, Thinking Machines or Sail Research?

Thinking Machines: Inkling up to 1M; Tinker 32K–256K. Sail Research: Varies by model.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.