Cerebras vs Wafer
Cerebras gets speed from custom silicon. Wafer gets it from agent-tuned software on NVIDIA and AMD GPUs. Different bets on the same goal.
By The Subconscious Team · Updated
Cerebras vs Wafer: key differences
Despite the name, Wafer does not build chips. It builds AI agents that tune inference stacks, searching across batching, decoding, quantization, engines and kernels, then keeps re-tuning a dedicated deployment as traffic changes. Wafer reports its tuned Qwen 3.5 397B running 2.8x faster than stock SGLang, and GLM 5.1 and DeepSeek V4 Pro each 2x faster than a vLLM baseline. Cerebras gets its speed from hardware, a wafer-scale chip that serves GPT-OSS 120B near 3,000 tokens per second, the fastest published figure of any public host.
The model mix differs. Wafer runs big open models like Qwen 3.5 397B and GLM 5.1, which suits coding agents that want large models at interactive speed. The Cerebras shared catalog is GPT-OSS 120B and Gemma 4 31B. Pricing differs too: Wafer Pass is a flat subscription from $10 a week across every hosted model, while Cerebras bills per token. Wafer is very young, and its speedups are self-reported against stock baselines. Cerebras is public on Nasdaq with OpenAI as its anchor customer.
What Cerebras and Wafer do
Cerebras
Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.
Example models: GPT-OSS 120B, Gemma 4 31B
Full Cerebras profileWafer
Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.
Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo
Full Wafer profileShould you choose Cerebras or Wafer?
Cerebras
Choose Cerebras for
- Maximum published output speed on GPT-OSS 120B
- Voice and streaming UIs
- Buyers who want a public company with a major anchor customer
Wafer
Choose Wafer for
- Large open models like Qwen 3.5 397B at interactive speed
- Flat-rate coding agent access from $10 a week
- Hedging GPU supply across NVIDIA and AMD
Cerebras vs Wafer at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GPT-OSS 120B, Gemma 4 31B | Qwen 3.5 397B Turbo, GLM 5.1 Turbo |
| Speed | ~3,000 tok/s on GPT-OSS 120B | 2–2.8x vs stock vLLM or SGLang |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) | Wafer Pass from $10 a week |
| Customization | Unknown | Agent-tuned dedicated deployments |
| Deployment | Shared API, dedicated, partners | Serverless pass, dedicated |
| Long context | Unknown | Varies by model |
Frequently asked questions
What is the difference between Cerebras and Wafer?
Cerebras gets speed from custom silicon. Wafer gets it from agent-tuned software on NVIDIA and AMD GPUs. Different bets on the same goal.
When should I choose Cerebras over Wafer?
Maximum published output speed on GPT-OSS 120B; Voice and streaming UIs; Buyers who want a public company with a major anchor customer.
When should I choose Wafer over Cerebras?
Large open models like Qwen 3.5 397B at interactive speed; Flat-rate coding agent access from $10 a week; Hedging GPU supply across NVIDIA and AMD.
Is Cerebras or Wafer cheaper?
Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.