We raised $5.1M for long-running agents.
vs

Crusoe vs Thinking Machines

Thinking Machines sells Tinker, a low-level post-training API, plus its own Inkling models. Crusoe is a production inference and GPU cloud with managed LoRA.

By The Subconscious Team · Updated

Crusoe vs Thinking Machines: key differences

Thinking Machines is not a general inference host. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so research teams write their own supervised or reinforcement learning loops while the lab runs distributed GPU work. It trains LoRA adapters on Qwen3.5, Nemotron 3, GLM-5.3, Kimi K2.6, DeepSeek-V3.1, gpt-oss and its own Inkling models. Billing splits into prefill, sample and train meters, with GPT-OSS-20B at $0.18, $0.45 and $0.40 per million. Crusoe's LoRA fine-tuning is a managed service attached to a serving stack, aimed at teams that want a tuned model in production rather than full control of the training loop.

Serving is where Crusoe pulls ahead. Its serverless API covers DeepSeek, GLM, Kimi, Gemma, gpt-oss and Nemotron from $0.05 in and $0.20 out, backed by self-serve and tailored dedicated deployments and a cluster-wide KV cache. Thinking Machines' beta serverless API serves only Inkling and Inkling-Small, with Inkling at $1.00 in and $4.05 out, and its docs scope checkpoint sampling to testing and low internal traffic. The Inkling models are the lab's draw: Apache 2.0, text, image and audio input, and up to 1M context. Tinker's training context runs 32K to 256K depending on the model.

What Crusoe and Thinking Machines do

Crusoe

Crusoe started in 2018 turning wasted natural gas into power for computing and has since become a vertically integrated AI infrastructure company: it sources energy, builds data centers and rents GPUs through Crusoe Cloud. It designed and built the Abilene, Texas campus behind the OpenAI and Oracle Stargate project, planned at 1.2 GW, and in March 2026 announced an adjacent 900 MW campus for Microsoft. On September 17, 2026 it closed the first part of a $3.9B Series F at a $30.9B post-money valuation, and it reports over 6 GW of contracted capacity. Crusoe Cloud lists GB200 NVL72, B200 and AMD MI355X by quote, with H100 at $3.90 and H200 at $4.29 per GPU-hour on demand.

Example models: DeepSeek V4 Pro, GLM 5.3, Kimi K2.6

Full Crusoe profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Crusoe or Thinking Machines?

Crusoe

Choose Crusoe for

  • Serving tuned open models to real users
  • Managed LoRA without writing training code
  • Dedicated endpoints with SLAs

Thinking Machines

Choose Thinking Machines for

  • Custom RL and SFT loops on open models
  • Post-training large MoE models like Kimi K2.6
  • Trying the multimodal Inkling models

Crusoe vs Thinking Machines at a glance

AttributeCrusoeThinking Machines
Model accessOpen weightsOpen weights
Flagship modelsDeepSeek V4, GLM 5.3, Kimi K2.6, Nemotron 3Inkling, Inkling-Small
SpeedUp to 9.9x faster TTFT vs vLLM (vendor claim)Unknown
Price$0.05–$1.74 in, $0.20–$4.40 out per 1MPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationServerless LoRA fine-tuningLoRA SFT and RL via Tinker
DeploymentServerless, self-serve and tailored dedicated, raw GPUsTraining API, beta serverless (Inkling only)
Long contextVaries by model; cluster-wide KV cacheInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Crusoe and Thinking Machines?

Thinking Machines sells Tinker, a low-level post-training API, plus its own Inkling models. Crusoe is a production inference and GPU cloud with managed LoRA.

When should I choose Crusoe over Thinking Machines?

Serving tuned open models to real users; Managed LoRA without writing training code; Dedicated endpoints with SLAs.

When should I choose Thinking Machines over Crusoe?

Custom RL and SFT loops on open models; Post-training large MoE models like Kimi K2.6; Trying the multimodal Inkling models.

Is Crusoe or Thinking Machines cheaper?

Crusoe: $0.05–$1.74 in, $0.20–$4.40 out per 1M. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Crusoe or Thinking Machines?

Crusoe: Varies by model; cluster-wide KV cache. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.