Nebius vs Thinking Machines
Nebius is a full AI cloud with 60+ hosted open models and raw GPUs. Thinking Machines trains open models through Tinker and serves little beyond Inkling.
By The Subconscious Team · Updated
Nebius vs Thinking Machines: key differences
Nebius covers more ground. Token Factory serves 60+ open models, including Qwen, DeepSeek, GLM, Kimi and GPT-OSS, from $0.06 per million input tokens, with dedicated endpoints under a 99.9% SLA and optional EU or US placement. Teams can upload a fine-tuned checkpoint and serve it at the same token pricing, and rent H100s or GB300 racks on the same account. Thinking Machines does one thing in depth. Tinker lets researchers write custom SFT or RL loops with forward_backward, optim_step, sample and save_state, while the lab handles distributed training on LoRA adapters. It trains Qwen3.5, GLM-5.3, Kimi K2.6, DeepSeek-V3.1, gpt-oss and its own Inkling models.
The two can pair well. Tinker produces the adapter, and a host like Nebius serves it with an SLA, since Tinker's OpenAI-compatible sampling endpoint is scoped to testing and low internal traffic. Nebius does not advertise an RL training API, so teams wanting that level of control would otherwise need to run their own training on Nebius GPUs. Thinking Machines has a unique asset in Inkling, a 975B MoE with 41B active and 1M context across text, image and audio, served in beta at $1.00 in and $4.05 out. Nebius has a $25 minimum first payment and no free trial. For regulated European production traffic, Nebius wins clearly.
What Nebius and Thinking Machines do
Nebius
Nebius is an Amsterdam-headquartered AI cloud and the strongest European alternative to the US hyperscalers. It sells raw NVIDIA GPU compute, from H100s at $2.15 an hour preemptible up to GB300 NVL72 racks, and it has begun adding Vera Rubin. Hyperscale buyers back it: a Microsoft capacity deal worth about $17.4B in September 2025, then a Meta agreement worth up to about $27B in March 2026.
Example models: DeepSeek V3, GPT-OSS
Full Nebius profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Nebius or Thinking Machines?
Nebius
Choose Nebius for
- EU data residency for production inference
- Serving uploaded fine-tunes under a 99.9% SLA
- Growing from tokens into raw GPU training
Thinking Machines
Choose Thinking Machines for
- Writing custom RL loops without running clusters
- LoRA post-training on Kimi K2.6 or Inkling
- Sampling checkpoints mid-training
Nebius vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights, 60+ models | Open weights |
| Flagship models | DeepSeek, Qwen, GLM, Kimi, GPT-OSS | Inkling, Inkling-Small |
| Speed | Among top hosts on throughput | Unknown |
| Price | From $0.06 per 1M input | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Serve uploaded fine-tunes | LoRA SFT and RL via Tinker |
| Deployment | Token Factory, dedicated, raw GPUs | Training API, beta serverless (Inkling only) |
| Long context | Varies by model | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Nebius and Thinking Machines?
Nebius is a full AI cloud with 60+ hosted open models and raw GPUs. Thinking Machines trains open models through Tinker and serves little beyond Inkling.
When should I choose Nebius over Thinking Machines?
EU data residency for production inference; Serving uploaded fine-tunes under a 99.9% SLA; Growing from tokens into raw GPU training.
When should I choose Thinking Machines over Nebius?
Writing custom RL loops without running clusters; LoRA post-training on Kimi K2.6 or Inkling; Sampling checkpoints mid-training.
Is Nebius or Thinking Machines cheaper?
Nebius: From $0.06 per 1M input. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Nebius or Thinking Machines?
Nebius: Varies by model. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Nebius
OpenAI vs Nebius
Anthropic vs Nebius
Google Vertex AI vs Nebius
Amazon Bedrock vs Nebius
Together AI vs Nebius
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.