We raised $5.1M for long-running agents.
vs

Cerebras vs Cloudflare Workers AI

On GPT-OSS 120B, Cerebras and Cloudflare Workers AI list the same $0.35 in and $0.75 out. Cerebras claims about 3,000 tokens per second; Cloudflare offers a far wider catalog.

By The Subconscious Team · Updated

Cerebras vs Cloudflare Workers AI: key differences

The cleanest head-to-head in this set: both list GPT-OSS 120B at $0.35 in and $0.75 out per million tokens. Cerebras's developer table puts it near 3,000 tokens per second on its wafer-scale chip, the fastest published figure of any public host. Cloudflare publishes no speed number for the same model, and its synchronous requests can queue for capacity. For the same money on the same weights, Cerebras is the speed pick. Cloudflare gets the edge on cost at low volume, with 10,000 free Neurons a day, and on protocol, since it adds a Responses endpoint for gpt-oss alongside OpenAI-compatible Chat Completions.

Breadth reverses the picture. Cerebras's shared catalog is just GPT-OSS 120B and Gemma 4 31B, and most other models mean a dedicated endpoint, a sales conversation or a partner like OpenRouter or AWS Marketplace. Cloudflare lists 50+ models, including DeepSeek V4 Pro with the full 1M context, GLM 5.3 and Kimi K2.7 Code, plus embeddings and bring-your-own LoRA on smaller models. Cerebras's speed also does little when an agent mostly waits on tools. Cerebras does offer a path to OpenAI's Ultrafast GPT-5.6 Sol preview. Pick Cerebras for streamed output where generation is the wait. Pick Cloudflare when an app needs several models and long context from one platform.

What Cerebras and Cloudflare Workers AI do

Cerebras

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

Example models: GPT-OSS 120B, Gemma 4 31B

Full Cerebras profile

Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

Example models: DeepSeek V4 Pro, GLM 5.3

Full Cloudflare Workers AI profile

Should you choose Cerebras or Cloudflare Workers AI?

Cerebras

Choose Cerebras for

  • Fastest published speed on GPT-OSS 120B
  • Live code autocomplete and voice streaming
  • Access to GPT-5.6 Sol Ultrafast

Cloudflare Workers AI

Choose Cloudflare Workers AI for

  • Many open models on one self-serve account
  • 1M context on DeepSeek V4 Pro
  • Tool-bound agents where raw speed matters less

Cerebras vs Cloudflare Workers AI at a glance

AttributeCerebrasCloudflare Workers AI
Model accessOpen weightsOpen weights
Flagship modelsGPT-OSS 120B, Gemma 4 31BDeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B
Speed~3,000 tok/s on GPT-OSS 120BUnknown
Price$0.35 in, $0.75 out (GPT-OSS 120B)$0.011 per 1K Neurons; 10K free daily
CustomizationUnknownBYO LoRA on small models (beta)
DeploymentShared API, dedicated, partnersServerless on Cloudflare network
Long contextUnknown1M on DeepSeek V4; 262K on Kimi

Frequently asked questions

What is the difference between Cerebras and Cloudflare Workers AI?

On GPT-OSS 120B, Cerebras and Cloudflare Workers AI list the same $0.35 in and $0.75 out. Cerebras claims about 3,000 tokens per second; Cloudflare offers a far wider catalog.

When should I choose Cerebras over Cloudflare Workers AI?

Fastest published speed on GPT-OSS 120B; Live code autocomplete and voice streaming; Access to GPT-5.6 Sol Ultrafast.

When should I choose Cloudflare Workers AI over Cerebras?

Many open models on one self-serve account; 1M context on DeepSeek V4 Pro; Tool-bound agents where raw speed matters less.

Is Cerebras or Cloudflare Workers AI cheaper?

Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.