We raised $5.1M for long-running agents.
vs

Nebius vs Thinking Machines

Nebius is a full AI cloud with 60+ hosted open models and raw GPUs. Thinking Machines trains open models through Tinker and serves little beyond Inkling.

By The Subconscious Team · Updated

Nebius vs Thinking Machines: key differences

Nebius covers more ground. Token Factory serves 60+ open models, including Qwen, DeepSeek, GLM, Kimi and GPT-OSS, from $0.06 per million input tokens, with dedicated endpoints under a 99.9% SLA and optional EU or US placement. Teams can upload a fine-tuned checkpoint and serve it at the same token pricing, and rent H100s or GB300 racks on the same account. Thinking Machines does one thing in depth. Tinker lets researchers write custom SFT or RL loops with forward_backward, optim_step, sample and save_state, while the lab handles distributed training on LoRA adapters. It trains Qwen3.5, GLM-5.3, Kimi K2.6, DeepSeek-V3.1, gpt-oss and its own Inkling models.

The two can pair well. Tinker produces the adapter, and a host like Nebius serves it with an SLA, since Tinker's OpenAI-compatible sampling endpoint is scoped to testing and low internal traffic. Nebius does not advertise an RL training API, so teams wanting that level of control would otherwise need to run their own training on Nebius GPUs. Thinking Machines has a unique asset in Inkling, a 975B MoE with 41B active and 1M context across text, image and audio, served in beta at $1.00 in and $4.05 out. Nebius has a $25 minimum first payment and no free trial. For regulated European production traffic, Nebius wins clearly.

What Nebius and Thinking Machines do

Nebius

Nebius is an Amsterdam-headquartered AI cloud and the strongest European alternative to the US hyperscalers. It sells raw NVIDIA GPU compute, from H100s at $2.15 an hour preemptible up to GB300 NVL72 racks, and it has begun adding Vera Rubin. Hyperscale buyers back it: a Microsoft capacity deal worth about $17.4B in September 2025, then a Meta agreement worth up to about $27B in March 2026.

Example models: DeepSeek V3, GPT-OSS

Full Nebius profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Nebius or Thinking Machines?

Nebius

Choose Nebius for

  • EU data residency for production inference
  • Serving uploaded fine-tunes under a 99.9% SLA
  • Growing from tokens into raw GPU training

Thinking Machines

Choose Thinking Machines for

  • Writing custom RL loops without running clusters
  • LoRA post-training on Kimi K2.6 or Inkling
  • Sampling checkpoints mid-training

Nebius vs Thinking Machines at a glance

AttributeNebiusThinking Machines
Model accessOpen weights, 60+ modelsOpen weights
Flagship modelsDeepSeek, Qwen, GLM, Kimi, GPT-OSSInkling, Inkling-Small
SpeedAmong top hosts on throughputUnknown
PriceFrom $0.06 per 1M inputPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationServe uploaded fine-tunesLoRA SFT and RL via Tinker
DeploymentToken Factory, dedicated, raw GPUsTraining API, beta serverless (Inkling only)
Long contextVaries by modelInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Nebius and Thinking Machines?

Nebius is a full AI cloud with 60+ hosted open models and raw GPUs. Thinking Machines trains open models through Tinker and serves little beyond Inkling.

When should I choose Nebius over Thinking Machines?

EU data residency for production inference; Serving uploaded fine-tunes under a 99.9% SLA; Growing from tokens into raw GPU training.

When should I choose Thinking Machines over Nebius?

Writing custom RL loops without running clusters; LoRA post-training on Kimi K2.6 or Inkling; Sampling checkpoints mid-training.

Is Nebius or Thinking Machines cheaper?

Nebius: From $0.06 per 1M input. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Nebius or Thinking Machines?

Nebius: Varies by model. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.