Cerebras vs SambaNova
Two custom-chip speed specialists. Cerebras posts the fastest public numbers on GPT-OSS; SambaNova targets fast decode on larger models it can hot swap.
By The Subconscious Team · Updated
Cerebras vs SambaNova: key differences
Cerebras and SambaNova both skip GPUs for their own silicon, and both sell speed. Cerebras builds a wafer-scale chip and lists GPT-OSS 120B near 3,000 tokens per second, the fastest published figure of any public host. SambaNova's RDU uses three tiers of memory to host very large models and hot swap between them in milliseconds, and SambaNova says its SN50 rack runs MiniMax M2.7 near 820 tokens per second in its fastest configuration. Each has a thin public catalog, and SambaNova's own profile notes that many of its numbers are vendor benchmarks on hardware still ramping.
The catalogs point at different users. Cerebras' shared API held GPT-OSS 120B and Gemma 4 31B as of August 2026. SambaCloud adds larger models such as MiniMax M2.7 and DeepSeek, which suits interactive coding agents on big open weights. Each also sells beyond the API. OpenAI rents about 750 MW of Cerebras capacity through 2028, and a new Gimlet Labs partnership pairs its wafers with GPUs. SambaNova sells air-cooled racks to neoclouds and splits prefill onto GPUs with decode on RDUs. For the fastest tokens on GPT-OSS, pick Cerebras. For speed on larger models, or agents that swap models often, pick SambaNova.
What Cerebras and SambaNova do
Cerebras
Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.
Example models: GPT-OSS 120B, Gemma 4 31B
Full Cerebras profileSambaNova
SambaNova designs its own inference chip, the Reconfigurable Dataflow Unit, and sells fast tokens on large open models through SambaCloud. The RDU maps the model graph onto the chip to cut trips to off-chip memory. A three-tier memory design of SRAM, HBM and bulk DRAM lets one system host very large models and hot swap between several of them in milliseconds. SambaCloud serves models like MiniMax M2.7, DeepSeek, Gemma 4 31B and GPT-OSS 120B, with speeds reported by Artificial Analysis.
Example models: MiniMax M2.7, GPT-OSS 120B
Full SambaNova profileShould you choose Cerebras or SambaNova?
Cerebras
Choose Cerebras for
- The fastest published output on GPT-OSS 120B
- Voice and live autocomplete on smaller open models
- A wafer-scale route to OpenAI's Ultrafast GPT-5.6 Sol
SambaNova
Choose SambaNova for
- Fast decode on larger models like MiniMax M2.7 and DeepSeek
- Agents that hot swap between several models
- Neoclouds buying air-cooled racks for a speed tier
Cerebras vs SambaNova at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GPT-OSS 120B, Gemma 4 31B | MiniMax M2.7, GPT-OSS 120B, DeepSeek |
| Speed | ~3,000 tok/s on GPT-OSS 120B | ~820 tok/s on MiniMax M2.7 (SN50) |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) | $0.22 in, $0.59 out (GPT-OSS 120B) |
| Customization | Unknown | Unknown |
| Deployment | Shared API, dedicated, partners | SambaCloud, racks for neoclouds |
| Long context | Unknown | Up to 192K (MiniMax M2.7) |
Frequently asked questions
What is the difference between Cerebras and SambaNova?
Two custom-chip speed specialists. Cerebras posts the fastest public numbers on GPT-OSS; SambaNova targets fast decode on larger models it can hot swap.
When should I choose Cerebras over SambaNova?
The fastest published output on GPT-OSS 120B; Voice and live autocomplete on smaller open models; A wafer-scale route to OpenAI's Ultrafast GPT-5.6 Sol.
When should I choose SambaNova over Cerebras?
Fast decode on larger models like MiniMax M2.7 and DeepSeek; Agents that hot swap between several models; Neoclouds buying air-cooled racks for a speed tier.
Is Cerebras or SambaNova cheaper?
Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). SambaNova: $0.22 in, $0.59 out (GPT-OSS 120B). The cheaper choice depends on the model and workload.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.