# Groq vs Cloudflare Workers AI

> Groq trades catalog size for LPU speed and tight tail latency. Cloudflare Workers AI trades speed guarantees for 50+ open models and up to 1M context.

Canonical: https://www.subconscious.dev/compare/groq-vs-cloudflare-workers-ai · By The Subconscious Team · Updated September 30, 2026

## How they compare

Groq serves open models on its own LPU and publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B, with median and tail latency staying close. Cloudflare runs GPUs in its network, publishes no speed figure, and warns that synchronous requests can queue for capacity. Both serve gpt-oss, and Cloudflare lists gpt-oss 120B at $0.35 in and $0.75 out. Groq's small-model prices sit near the market floor, with cache and Batch discounts that stack. Context is a clear divide. Groq caps around 131K tokens, while Cloudflare offers the full 1M on DeepSeek V4 and 262K on Kimi, which matters for agents carrying long histories.

Catalog direction differs too. Groq's list is narrowing, with Llama 3.3 70B and Llama 3.1 8B shut down on August 16, 2026. Cloudflare's has grown since March 2026 to include GLM 5.3, DeepSeek V4 Pro and Kimi K2.7 Code. Groq adds Whisper for speech to text and Groq Compound, with built-in search and code execution. Cloudflare adds embeddings, bring-your-own LoRA on smaller models and AI Gateway for fallbacks. Groq hosts no fine-tuned models, and its long-term outlook is uncertain since NVIDIA hired most of its engineers. For voice agents and strict latency SLAs, Groq wins. For long-context agents and broad model choice on one platform, Cloudflare wins.

## What each one does

### Groq

Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.

### Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

## Which is best, and when

### Choose Groq for

- Voice agents pairing Whisper with fast replies
- Strict SLAs judged on tail latency
- Tight agent loops on GPT-OSS

### Choose Cloudflare Workers AI for

- Agents that need more than 131K context
- Frontier-scale open models like DeepSeek V4 Pro
- LoRA adapters on smaller models

## At a glance

| Attribute | Groq | Cloudflare Workers AI |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GPT-OSS 120B, Qwen 3.6 27B | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B |
| Speed | 500–1,000 tok/s | - |
| Price | Near the floor on small models | $0.011 per 1K Neurons; 10K free daily |
| Customization | No fine-tuned model hosting | BYO LoRA on small models (beta) |
| Deployment | GroqCloud API | Serverless on Cloudflare network |
| Long context | Around 131K max | 1M on DeepSeek V4; 262K on Kimi |

## FAQ

### What is the difference between Groq and Cloudflare Workers AI?

Groq trades catalog size for LPU speed and tight tail latency. Cloudflare Workers AI trades speed guarantees for 50+ open models and up to 1M context.

### When should I choose Groq over Cloudflare Workers AI?

Voice agents pairing Whisper with fast replies; Strict SLAs judged on tail latency; Tight agent loops on GPT-OSS.

### When should I choose Cloudflare Workers AI over Groq?

Agents that need more than 131K context; Frontier-scale open models like DeepSeek V4 Pro; LoRA adapters on smaller models.

### Is Groq or Cloudflare Workers AI cheaper?

Groq: Near the floor on small models. Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. The cheaper choice depends on the model and workload.

### Which has more context, Groq or Cloudflare Workers AI?

Groq: Around 131K max. Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Groq](https://www.subconscious.dev/compare/subconscious-vs-groq.md), [Subconscious vs Cloudflare Workers AI](https://www.subconscious.dev/compare/subconscious-vs-cloudflare-workers-ai.md).

Full profiles: [Groq](https://www.subconscious.dev/providers/groq.md), [Cloudflare Workers AI](https://www.subconscious.dev/providers/cloudflare-workers-ai.md).
