# Cerebras vs Cloudflare Workers AI

> On GPT-OSS 120B, Cerebras and Cloudflare Workers AI list the same $0.35 in and $0.75 out. Cerebras claims about 3,000 tokens per second; Cloudflare offers a far wider catalog.

Canonical: https://www.subconscious.dev/compare/cerebras-vs-cloudflare-workers-ai · By The Subconscious Team · Updated September 30, 2026

## How they compare

The cleanest head-to-head in this set: both list GPT-OSS 120B at $0.35 in and $0.75 out per million tokens. Cerebras's developer table puts it near 3,000 tokens per second on its wafer-scale chip, the fastest published figure of any public host. Cloudflare publishes no speed number for the same model, and its synchronous requests can queue for capacity. For the same money on the same weights, Cerebras is the speed pick. Cloudflare gets the edge on cost at low volume, with 10,000 free Neurons a day, and on protocol, since it adds a Responses endpoint for gpt-oss alongside OpenAI-compatible Chat Completions.

Breadth reverses the picture. Cerebras's shared catalog is just GPT-OSS 120B and Gemma 4 31B, and most other models mean a dedicated endpoint, a sales conversation or a partner like OpenRouter or AWS Marketplace. Cloudflare lists 50+ models, including DeepSeek V4 Pro with the full 1M context, GLM 5.3 and Kimi K2.7 Code, plus embeddings and bring-your-own LoRA on smaller models. Cerebras's speed also does little when an agent mostly waits on tools. Cerebras does offer a path to OpenAI's Ultrafast GPT-5.6 Sol preview. Pick Cerebras for streamed output where generation is the wait. Pick Cloudflare when an app needs several models and long context from one platform.

## What each one does

### Cerebras

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

### Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

## Which is best, and when

### Choose Cerebras for

- Fastest published speed on GPT-OSS 120B
- Live code autocomplete and voice streaming
- Access to GPT-5.6 Sol Ultrafast

### Choose Cloudflare Workers AI for

- Many open models on one self-serve account
- 1M context on DeepSeek V4 Pro
- Tool-bound agents where raw speed matters less

## At a glance

| Attribute | Cerebras | Cloudflare Workers AI |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GPT-OSS 120B, Gemma 4 31B | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B |
| Speed | ~3,000 tok/s on GPT-OSS 120B | - |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) | $0.011 per 1K Neurons; 10K free daily |
| Customization | - | BYO LoRA on small models (beta) |
| Deployment | Shared API, dedicated, partners | Serverless on Cloudflare network |
| Long context | - | 1M on DeepSeek V4; 262K on Kimi |

## FAQ

### What is the difference between Cerebras and Cloudflare Workers AI?

On GPT-OSS 120B, Cerebras and Cloudflare Workers AI list the same $0.35 in and $0.75 out. Cerebras claims about 3,000 tokens per second; Cloudflare offers a far wider catalog.

### When should I choose Cerebras over Cloudflare Workers AI?

Fastest published speed on GPT-OSS 120B; Live code autocomplete and voice streaming; Access to GPT-5.6 Sol Ultrafast.

### When should I choose Cloudflare Workers AI over Cerebras?

Many open models on one self-serve account; 1M context on DeepSeek V4 Pro; Tool-bound agents where raw speed matters less.

### Is Cerebras or Cloudflare Workers AI cheaper?

Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Cerebras](https://www.subconscious.dev/compare/subconscious-vs-cerebras.md), [Subconscious vs Cloudflare Workers AI](https://www.subconscious.dev/compare/subconscious-vs-cloudflare-workers-ai.md).

Full profiles: [Cerebras](https://www.subconscious.dev/providers/cerebras.md), [Cloudflare Workers AI](https://www.subconscious.dev/providers/cloudflare-workers-ai.md).
