Novita AI vs Thinking Machines
Novita AI is a cheap, broad inference cloud with 200+ models. Thinking Machines is a post-training API with a narrow beta serving layer for Inkling.
By The Subconscious Team · Updated
Novita AI vs Thinking Machines: key differences
Novita wins on breadth and price. Its serverless API covers 200+ open models across LLMs, image, video and speech, with LLM prices from $0.02 per million tokens and batch at 50% off. DeepSeek V4 Pro runs with its full 1M context. Dedicated endpoints serve any Hugging Face model with hot-swappable LoRA adapters under a 99.5% SLA, and a GPU cloud and Firecracker agent sandbox share the bill. Thinking Machines charges more to serve, with Inkling at $1.00 in and $4.05 out, and its serverless API covers only Inkling and Inkling-Small. Its real product is Tinker, where teams run custom SFT and RL loops on open models.
That makes a common split: train on Tinker, serve elsewhere. Tinker handles distributed LoRA training on large MoE models like Kimi K2.6, GLM-5.3 and Inkling, which is hard to do in-house, and teams keep full control of the loss and sampling logic. Novita's hot-swappable LoRA endpoints are one place such an adapter could land, if the base model is on Novita's list. Novita does not offer a comparable training API. Buyers should weigh Novita's looser serverless SLAs, Discord-based support and lack of public SOC 2. Thinking Machines has deep funding and a large Nvidia capacity deal, but its checkpoint endpoint is not meant for user-facing traffic.
What Novita AI and Thinking Machines do
Novita AI
Novita AI is a San Francisco inference cloud founded in late 2023 by Frank Lewis and Junyu Huang, and it competes on price and breadth. Its serverless API covers 200+ open models across LLMs, image, video, speech, voice cloning and embeddings, with LLM prices starting at $0.02 per million tokens. The API speaks both OpenAI and Anthropic formats. It became an official Hugging Face Inference Partner in April 2026 and was the day-zero launch partner for Google's Gemma 4.
Example models: DeepSeek V4 Pro, Gemma 4
Full Novita AI profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Novita AI or Thinking Machines?
Novita AI
Choose Novita AI for
- Cost-first inference across 200+ models
- Serving LoRA adapters on dedicated endpoints
- Models, GPUs and sandboxes on one bill
Thinking Machines
Choose Thinking Machines for
- Custom RL post-training on open weights
- Training adapters on large MoE bases
- Evaluating Inkling's text, image and audio input
Novita AI vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | DeepSeek V4 Pro, Gemma 4 | Inkling, Inkling-Small |
| Speed | ~36 tok/s on DeepSeek V4 Pro | Unknown |
| Price | From $0.02 per 1M; batch 50% off | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Hot-swappable LoRA adapters | LoRA SFT and RL via Tinker |
| Deployment | Serverless, GPU cloud, dedicated | Training API, beta serverless (Inkling only) |
| Long context | Full 1M on DeepSeek V4 Pro | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Novita AI and Thinking Machines?
Novita AI is a cheap, broad inference cloud with 200+ models. Thinking Machines is a post-training API with a narrow beta serving layer for Inkling.
When should I choose Novita AI over Thinking Machines?
Cost-first inference across 200+ models; Serving LoRA adapters on dedicated endpoints; Models, GPUs and sandboxes on one bill.
When should I choose Thinking Machines over Novita AI?
Custom RL post-training on open weights; Training adapters on large MoE bases; Evaluating Inkling's text, image and audio input.
Is Novita AI or Thinking Machines cheaper?
Novita AI: From $0.02 per 1M; batch 50% off. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Novita AI or Thinking Machines?
Novita AI: Full 1M on DeepSeek V4 Pro. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Novita AI
OpenAI vs Novita AI
Anthropic vs Novita AI
Google Vertex AI vs Novita AI
Amazon Bedrock vs Novita AI
Together AI vs Novita AI
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.