Cerebras vs Thinking Machines
Cerebras runs a tiny catalog at wafer-scale speed. Thinking Machines trains open models and ships Inkling. Speed at inference versus control at training time.
By The Subconscious Team · Updated
Cerebras vs Thinking Machines: key differences
Cerebras is the fastest public inference host on the models it serves, with GPT-OSS 120B listed near 3,000 tokens per second at $0.35 in and $0.75 out. Its shared catalog is just GPT-OSS 120B and Gemma 4 31B, with more families on dedicated endpoints and partners like OpenRouter and AWS Marketplace, and it powers OpenAI's Ultrafast GPT-5.6 Sol preview. Cerebras does not pitch customization. Thinking Machines publishes no speed figures at all. Its product is Tinker, which lets teams write their own SFT or RL loops with LoRA adapters on open models, including gpt-oss, Kimi K2.6, GLM-5.3 and its own Inkling family, billed by prefill, sample and train tokens.
The overlap is mostly around gpt-oss. A team could post-train gpt-oss on Tinker, but Cerebras' shared API would not serve that fine-tune; custom models there mean a dedicated endpoint and a sales conversation. Thinking Machines' own serving is limited to a beta serverless API for Inkling, at $1.00 in and $4.05 out with up to 1M context and native image and audio input, plus a checkpoint endpoint scoped to testing. Cerebras wins where generation time is the wait, like voice and live autocomplete. Thinking Machines wins where the base model is not good enough and the answer is training, not faster decoding.
What Cerebras and Thinking Machines do
Cerebras
Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.
Example models: GPT-OSS 120B, Gemma 4 31B
Full Cerebras profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Cerebras or Thinking Machines?
Cerebras
Choose Cerebras for
- Maximum tokens per second on GPT-OSS 120B
- Live autocomplete and streaming UIs
- Wafer-scale access to GPT-5.6 Sol Ultrafast
Thinking Machines
Choose Thinking Machines for
- Fine-tuning gpt-oss and larger MoE bases
- Multimodal Inkling models with audio input
- Custom RL research without cluster management
Cerebras vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GPT-OSS 120B, Gemma 4 31B | Inkling, Inkling-Small |
| Speed | ~3,000 tok/s on GPT-OSS 120B | Unknown |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Unknown | LoRA SFT and RL via Tinker |
| Deployment | Shared API, dedicated, partners | Training API, beta serverless (Inkling only) |
| Long context | Unknown | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Cerebras and Thinking Machines?
Cerebras runs a tiny catalog at wafer-scale speed. Thinking Machines trains open models and ships Inkling. Speed at inference versus control at training time.
When should I choose Cerebras over Thinking Machines?
Maximum tokens per second on GPT-OSS 120B; Live autocomplete and streaming UIs; Wafer-scale access to GPT-5.6 Sol Ultrafast.
When should I choose Thinking Machines over Cerebras?
Fine-tuning gpt-oss and larger MoE bases; Multimodal Inkling models with audio input; Custom RL research without cluster management.
Is Cerebras or Thinking Machines cheaper?
Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Related comparisons
Subconscious vs Cerebras
OpenAI vs Cerebras
Anthropic vs Cerebras
Google Vertex AI vs Cerebras
Amazon Bedrock vs Cerebras
Together AI vs Cerebras
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.