# Hugging Face Inference Providers vs Thinking Machines

> Thinking Machines sells Tinker for post-training open models, plus its own Inkling models. Hugging Face routes inference and offers no fine-tuning.

Canonical: https://www.subconscious.dev/compare/hugging-face-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

These products sit at opposite ends of the model lifecycle. Tinker, Thinking Machines' main developer product, exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while the lab runs distributed GPU work with LoRA adapters. It trains Qwen3.5, Nemotron 3, GLM-5.3, Kimi K2.6, DeepSeek-V3.1, gpt-oss and the lab's own Apache 2.0 Inkling models. Billing is per million tokens across prefill, sample and train meters. Hugging Face Inference Providers does no training at all. It routes 132 chat models across 17 partner hosts at their rates with no markup.

For serving, the router is the general-purpose option. Thinking Machines' beta serverless API covers only Inkling, at $1.00 in and $4.05 out, and Inkling-Small, and its OpenAI-compatible checkpoint endpoint is scoped to testing and low internal traffic. Hugging Face picks the fastest or cheapest host per model, fails over when a provider is down, and reaches up to 1M context depending on the provider. Inkling has its own draw: a 975B mixture-of-experts model with 41B active, text, image and audio input, and 1M context. Tinker's training context runs 32K to 256K. A team could train on Tinker and serve elsewhere, possibly on Hugging Face's dedicated Inference Endpoints, since the router only serves partner-hosted models.

## What each one does

### Hugging Face Inference Providers

Inference Providers is a router run by Hugging Face that sits in front of partner inference clouds. The current partner list covers Baseten, Cerebras, Cohere, DeepInfra, fal, Featherless AI, Fireworks, Groq, Novita, Nscale, OVHcloud, Public AI, Replicate, Scaleway, Together, WaveSpeedAI and Z.ai, plus Hugging Face's own HF Inference, which now mostly serves CPU workloads like embeddings and classification. Chat traffic goes through an OpenAI-compatible endpoint at router.huggingface.co/v1, and the Python and JavaScript clients add text-to-image, video, speech and embeddings. The router lists 132 chat models today, from GLM 5.3 and Kimi K3 to gpt-oss-120b on eleven providers.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Hugging Face Inference Providers for

- Serving popular open models in production
- Routing across hosts with failover
- Prototyping before committing to a model

### Choose Thinking Machines for

- Custom SFT or RL loops without running clusters
- LoRA fine-tuning of large MoE models like Kimi K2.6
- Evaluating the multimodal Inkling models

## At a glance

| Attribute | Hugging Face Inference Providers | Thinking Machines |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GLM 5.3, Kimi K3, DeepSeek V4.1 Flash | Inkling, Inkling-Small |
| Speed | Routes to fastest provider by default | - |
| Price | Provider rates, no markup | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | N/A | LoRA SFT and RL via Tinker |
| Deployment | Serverless router; dedicated Endpoints | Training API, beta serverless (Inkling only) |
| Long context | Up to 1M, provider-dependent | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Hugging Face Inference Providers and Thinking Machines?

Thinking Machines sells Tinker for post-training open models, plus its own Inkling models. Hugging Face routes inference and offers no fine-tuning.

### When should I choose Hugging Face Inference Providers over Thinking Machines?

Serving popular open models in production; Routing across hosts with failover; Prototyping before committing to a model.

### When should I choose Thinking Machines over Hugging Face Inference Providers?

Custom SFT or RL loops without running clusters; LoRA fine-tuning of large MoE models like Kimi K2.6; Evaluating the multimodal Inkling models.

### Is Hugging Face Inference Providers or Thinking Machines cheaper?

Hugging Face Inference Providers: Provider rates, no markup. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Hugging Face Inference Providers or Thinking Machines?

Hugging Face Inference Providers: Up to 1M, provider-dependent. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/subconscious-vs-hugging-face.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Hugging Face Inference Providers](https://www.subconscious.dev/providers/hugging-face.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
