# Novita AI vs Thinking Machines

> Novita AI is a cheap, broad inference cloud with 200+ models. Thinking Machines is a post-training API with a narrow beta serving layer for Inkling.

Canonical: https://www.subconscious.dev/compare/novita-ai-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Novita wins on breadth and price. Its serverless API covers 200+ open models across LLMs, image, video and speech, with LLM prices from $0.02 per million tokens and batch at 50% off. DeepSeek V4 Pro runs with its full 1M context. Dedicated endpoints serve any Hugging Face model with hot-swappable LoRA adapters under a 99.5% SLA, and a GPU cloud and Firecracker agent sandbox share the bill. Thinking Machines charges more to serve, with Inkling at $1.00 in and $4.05 out, and its serverless API covers only Inkling and Inkling-Small. Its real product is Tinker, where teams run custom SFT and RL loops on open models.

That makes a common split: train on Tinker, serve elsewhere. Tinker handles distributed LoRA training on large MoE models like Kimi K2.6, GLM-5.3 and Inkling, which is hard to do in-house, and teams keep full control of the loss and sampling logic. Novita's hot-swappable LoRA endpoints are one place such an adapter could land, if the base model is on Novita's list. Novita does not offer a comparable training API. Buyers should weigh Novita's looser serverless SLAs, Discord-based support and lack of public SOC 2. Thinking Machines has deep funding and a large Nvidia capacity deal, but its checkpoint endpoint is not meant for user-facing traffic.

## What each one does

### Novita AI

Novita AI is a San Francisco inference cloud founded in late 2023 by Frank Lewis and Junyu Huang, and it competes on price and breadth. Its serverless API covers 200+ open models across LLMs, image, video, speech, voice cloning and embeddings, with LLM prices starting at $0.02 per million tokens. The API speaks both OpenAI and Anthropic formats. It became an official Hugging Face Inference Partner in April 2026 and was the day-zero launch partner for Google's Gemma 4.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Novita AI for

- Cost-first inference across 200+ models
- Serving LoRA adapters on dedicated endpoints
- Models, GPUs and sandboxes on one bill

### Choose Thinking Machines for

- Custom RL post-training on open weights
- Training adapters on large MoE bases
- Evaluating Inkling's text, image and audio input

## At a glance

| Attribute | Novita AI | Thinking Machines |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | DeepSeek V4 Pro, Gemma 4 | Inkling, Inkling-Small |
| Speed | ~36 tok/s on DeepSeek V4 Pro | - |
| Price | From $0.02 per 1M; batch 50% off | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Hot-swappable LoRA adapters | LoRA SFT and RL via Tinker |
| Deployment | Serverless, GPU cloud, dedicated | Training API, beta serverless (Inkling only) |
| Long context | Full 1M on DeepSeek V4 Pro | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Novita AI and Thinking Machines?

Novita AI is a cheap, broad inference cloud with 200+ models. Thinking Machines is a post-training API with a narrow beta serving layer for Inkling.

### When should I choose Novita AI over Thinking Machines?

Cost-first inference across 200+ models; Serving LoRA adapters on dedicated endpoints; Models, GPUs and sandboxes on one bill.

### When should I choose Thinking Machines over Novita AI?

Custom RL post-training on open weights; Training adapters on large MoE bases; Evaluating Inkling's text, image and audio input.

### Is Novita AI or Thinking Machines cheaper?

Novita AI: From $0.02 per 1M; batch 50% off. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Novita AI or Thinking Machines?

Novita AI: Full 1M on DeepSeek V4 Pro. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Novita AI](https://www.subconscious.dev/compare/subconscious-vs-novita-ai.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Novita AI](https://www.subconscious.dev/providers/novita-ai.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
