# Cerebras vs Thinking Machines

> Cerebras runs a tiny catalog at wafer-scale speed. Thinking Machines trains open models and ships Inkling. Speed at inference versus control at training time.

Canonical: https://www.subconscious.dev/compare/cerebras-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Cerebras is the fastest public inference host on the models it serves, with GPT-OSS 120B listed near 3,000 tokens per second at $0.35 in and $0.75 out. Its shared catalog is just GPT-OSS 120B and Gemma 4 31B, with more families on dedicated endpoints and partners like OpenRouter and AWS Marketplace, and it powers OpenAI's Ultrafast GPT-5.6 Sol preview. Cerebras does not pitch customization. Thinking Machines publishes no speed figures at all. Its product is Tinker, which lets teams write their own SFT or RL loops with LoRA adapters on open models, including gpt-oss, Kimi K2.6, GLM-5.3 and its own Inkling family, billed by prefill, sample and train tokens.

The overlap is mostly around gpt-oss. A team could post-train gpt-oss on Tinker, but Cerebras' shared API would not serve that fine-tune; custom models there mean a dedicated endpoint and a sales conversation. Thinking Machines' own serving is limited to a beta serverless API for Inkling, at $1.00 in and $4.05 out with up to 1M context and native image and audio input, plus a checkpoint endpoint scoped to testing. Cerebras wins where generation time is the wait, like voice and live autocomplete. Thinking Machines wins where the base model is not good enough and the answer is training, not faster decoding.

## What each one does

### Cerebras

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Cerebras for

- Maximum tokens per second on GPT-OSS 120B
- Live autocomplete and streaming UIs
- Wafer-scale access to GPT-5.6 Sol Ultrafast

### Choose Thinking Machines for

- Fine-tuning gpt-oss and larger MoE bases
- Multimodal Inkling models with audio input
- Custom RL research without cluster management

## At a glance

| Attribute | Cerebras | Thinking Machines |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GPT-OSS 120B, Gemma 4 31B | Inkling, Inkling-Small |
| Speed | ~3,000 tok/s on GPT-OSS 120B | - |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | - | LoRA SFT and RL via Tinker |
| Deployment | Shared API, dedicated, partners | Training API, beta serverless (Inkling only) |
| Long context | - | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Cerebras and Thinking Machines?

Cerebras runs a tiny catalog at wafer-scale speed. Thinking Machines trains open models and ships Inkling. Speed at inference versus control at training time.

### When should I choose Cerebras over Thinking Machines?

Maximum tokens per second on GPT-OSS 120B; Live autocomplete and streaming UIs; Wafer-scale access to GPT-5.6 Sol Ultrafast.

### When should I choose Thinking Machines over Cerebras?

Fine-tuning gpt-oss and larger MoE bases; Multimodal Inkling models with audio input; Custom RL research without cluster management.

### Is Cerebras or Thinking Machines cheaper?

Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Cerebras](https://www.subconscious.dev/compare/subconscious-vs-cerebras.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Cerebras](https://www.subconscious.dev/providers/cerebras.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
