Cerebras vs Cloudflare Workers AI
On GPT-OSS 120B, Cerebras and Cloudflare Workers AI list the same $0.35 in and $0.75 out. Cerebras claims about 3,000 tokens per second; Cloudflare offers a far wider catalog.
By The Subconscious Team · Updated
Cerebras vs Cloudflare Workers AI: key differences
The cleanest head-to-head in this set: both list GPT-OSS 120B at $0.35 in and $0.75 out per million tokens. Cerebras's developer table puts it near 3,000 tokens per second on its wafer-scale chip, the fastest published figure of any public host. Cloudflare publishes no speed number for the same model, and its synchronous requests can queue for capacity. For the same money on the same weights, Cerebras is the speed pick. Cloudflare gets the edge on cost at low volume, with 10,000 free Neurons a day, and on protocol, since it adds a Responses endpoint for gpt-oss alongside OpenAI-compatible Chat Completions.
Breadth reverses the picture. Cerebras's shared catalog is just GPT-OSS 120B and Gemma 4 31B, and most other models mean a dedicated endpoint, a sales conversation or a partner like OpenRouter or AWS Marketplace. Cloudflare lists 50+ models, including DeepSeek V4 Pro with the full 1M context, GLM 5.3 and Kimi K2.7 Code, plus embeddings and bring-your-own LoRA on smaller models. Cerebras's speed also does little when an agent mostly waits on tools. Cerebras does offer a path to OpenAI's Ultrafast GPT-5.6 Sol preview. Pick Cerebras for streamed output where generation is the wait. Pick Cloudflare when an app needs several models and long context from one platform.
What Cerebras and Cloudflare Workers AI do
Cerebras
Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.
Example models: GPT-OSS 120B, Gemma 4 31B
Full Cerebras profileCloudflare Workers AI
Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.
Example models: DeepSeek V4 Pro, GLM 5.3
Full Cloudflare Workers AI profileShould you choose Cerebras or Cloudflare Workers AI?
Cerebras
Choose Cerebras for
- Fastest published speed on GPT-OSS 120B
- Live code autocomplete and voice streaming
- Access to GPT-5.6 Sol Ultrafast
Cloudflare Workers AI
Choose Cloudflare Workers AI for
- Many open models on one self-serve account
- 1M context on DeepSeek V4 Pro
- Tool-bound agents where raw speed matters less
Cerebras vs Cloudflare Workers AI at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GPT-OSS 120B, Gemma 4 31B | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B |
| Speed | ~3,000 tok/s on GPT-OSS 120B | Unknown |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) | $0.011 per 1K Neurons; 10K free daily |
| Customization | Unknown | BYO LoRA on small models (beta) |
| Deployment | Shared API, dedicated, partners | Serverless on Cloudflare network |
| Long context | Unknown | 1M on DeepSeek V4; 262K on Kimi |
Frequently asked questions
What is the difference between Cerebras and Cloudflare Workers AI?
On GPT-OSS 120B, Cerebras and Cloudflare Workers AI list the same $0.35 in and $0.75 out. Cerebras claims about 3,000 tokens per second; Cloudflare offers a far wider catalog.
When should I choose Cerebras over Cloudflare Workers AI?
Fastest published speed on GPT-OSS 120B; Live code autocomplete and voice streaming; Access to GPT-5.6 Sol Ultrafast.
When should I choose Cloudflare Workers AI over Cerebras?
Many open models on one self-serve account; 1M context on DeepSeek V4 Pro; Tool-bound agents where raw speed matters less.
Is Cerebras or Cloudflare Workers AI cheaper?
Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. The cheaper choice depends on the model and workload.
Related comparisons
Subconscious vs Cerebras
OpenAI vs Cerebras
Anthropic vs Cerebras
Google Vertex AI vs Cerebras
Amazon Bedrock vs Cerebras
Together AI vs Cerebras
Subconscious vs Cloudflare Workers AI
OpenAI vs Cloudflare Workers AI
Anthropic vs Cloudflare Workers AI
Google Vertex AI vs Cloudflare Workers AI
Amazon Bedrock vs Cloudflare Workers AI
Together AI vs Cloudflare Workers AI
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.