We raised $5.1M for long-running agents.
vs

Novita AI vs Thinking Machines

Novita AI is a cheap, broad inference cloud with 200+ models. Thinking Machines is a post-training API with a narrow beta serving layer for Inkling.

By The Subconscious Team · Updated

Novita AI vs Thinking Machines: key differences

Novita wins on breadth and price. Its serverless API covers 200+ open models across LLMs, image, video and speech, with LLM prices from $0.02 per million tokens and batch at 50% off. DeepSeek V4 Pro runs with its full 1M context. Dedicated endpoints serve any Hugging Face model with hot-swappable LoRA adapters under a 99.5% SLA, and a GPU cloud and Firecracker agent sandbox share the bill. Thinking Machines charges more to serve, with Inkling at $1.00 in and $4.05 out, and its serverless API covers only Inkling and Inkling-Small. Its real product is Tinker, where teams run custom SFT and RL loops on open models.

That makes a common split: train on Tinker, serve elsewhere. Tinker handles distributed LoRA training on large MoE models like Kimi K2.6, GLM-5.3 and Inkling, which is hard to do in-house, and teams keep full control of the loss and sampling logic. Novita's hot-swappable LoRA endpoints are one place such an adapter could land, if the base model is on Novita's list. Novita does not offer a comparable training API. Buyers should weigh Novita's looser serverless SLAs, Discord-based support and lack of public SOC 2. Thinking Machines has deep funding and a large Nvidia capacity deal, but its checkpoint endpoint is not meant for user-facing traffic.

What Novita AI and Thinking Machines do

Novita AI

Novita AI is a San Francisco inference cloud founded in late 2023 by Frank Lewis and Junyu Huang, and it competes on price and breadth. Its serverless API covers 200+ open models across LLMs, image, video, speech, voice cloning and embeddings, with LLM prices starting at $0.02 per million tokens. The API speaks both OpenAI and Anthropic formats. It became an official Hugging Face Inference Partner in April 2026 and was the day-zero launch partner for Google's Gemma 4.

Example models: DeepSeek V4 Pro, Gemma 4

Full Novita AI profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Novita AI or Thinking Machines?

Novita AI

Choose Novita AI for

  • Cost-first inference across 200+ models
  • Serving LoRA adapters on dedicated endpoints
  • Models, GPUs and sandboxes on one bill

Thinking Machines

Choose Thinking Machines for

  • Custom RL post-training on open weights
  • Training adapters on large MoE bases
  • Evaluating Inkling's text, image and audio input

Novita AI vs Thinking Machines at a glance

AttributeNovita AIThinking Machines
Model accessOpen weightsOpen weights
Flagship modelsDeepSeek V4 Pro, Gemma 4Inkling, Inkling-Small
Speed~36 tok/s on DeepSeek V4 ProUnknown
PriceFrom $0.02 per 1M; batch 50% offPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationHot-swappable LoRA adaptersLoRA SFT and RL via Tinker
DeploymentServerless, GPU cloud, dedicatedTraining API, beta serverless (Inkling only)
Long contextFull 1M on DeepSeek V4 ProInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Novita AI and Thinking Machines?

Novita AI is a cheap, broad inference cloud with 200+ models. Thinking Machines is a post-training API with a narrow beta serving layer for Inkling.

When should I choose Novita AI over Thinking Machines?

Cost-first inference across 200+ models; Serving LoRA adapters on dedicated endpoints; Models, GPUs and sandboxes on one bill.

When should I choose Thinking Machines over Novita AI?

Custom RL post-training on open weights; Training adapters on large MoE bases; Evaluating Inkling's text, image and audio input.

Is Novita AI or Thinking Machines cheaper?

Novita AI: From $0.02 per 1M; batch 50% off. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Novita AI or Thinking Machines?

Novita AI: Full 1M on DeepSeek V4 Pro. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.