We raised $5.1M for long-running agents.
vs

Groq vs Cloudflare Workers AI

Groq trades catalog size for LPU speed and tight tail latency. Cloudflare Workers AI trades speed guarantees for 50+ open models and up to 1M context.

By The Subconscious Team · Updated

Groq vs Cloudflare Workers AI: key differences

Groq serves open models on its own LPU and publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B, with median and tail latency staying close. Cloudflare runs GPUs in its network, publishes no speed figure, and warns that synchronous requests can queue for capacity. Both serve gpt-oss, and Cloudflare lists gpt-oss 120B at $0.35 in and $0.75 out. Groq's small-model prices sit near the market floor, with cache and Batch discounts that stack. Context is a clear divide. Groq caps around 131K tokens, while Cloudflare offers the full 1M on DeepSeek V4 and 262K on Kimi, which matters for agents carrying long histories.

Catalog direction differs too. Groq's list is narrowing, with Llama 3.3 70B and Llama 3.1 8B shut down on August 16, 2026. Cloudflare's has grown since March 2026 to include GLM 5.3, DeepSeek V4 Pro and Kimi K2.7 Code. Groq adds Whisper for speech to text and Groq Compound, with built-in search and code execution. Cloudflare adds embeddings, bring-your-own LoRA on smaller models and AI Gateway for fallbacks. Groq hosts no fine-tuned models, and its long-term outlook is uncertain since NVIDIA hired most of its engineers. For voice agents and strict latency SLAs, Groq wins. For long-context agents and broad model choice on one platform, Cloudflare wins.

What Groq and Cloudflare Workers AI do

Groq

Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.

Example models: GPT-OSS 120B, Qwen 3.6 27B

Full Groq profile

Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

Example models: DeepSeek V4 Pro, GLM 5.3

Full Cloudflare Workers AI profile

Should you choose Groq or Cloudflare Workers AI?

Groq

Choose Groq for

  • Voice agents pairing Whisper with fast replies
  • Strict SLAs judged on tail latency
  • Tight agent loops on GPT-OSS

Cloudflare Workers AI

Choose Cloudflare Workers AI for

  • Agents that need more than 131K context
  • Frontier-scale open models like DeepSeek V4 Pro
  • LoRA adapters on smaller models

Groq vs Cloudflare Workers AI at a glance

AttributeGroqCloudflare Workers AI
Model accessOpen weightsOpen weights
Flagship modelsGPT-OSS 120B, Qwen 3.6 27BDeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B
Speed500–1,000 tok/sUnknown
PriceNear the floor on small models$0.011 per 1K Neurons; 10K free daily
CustomizationNo fine-tuned model hostingBYO LoRA on small models (beta)
DeploymentGroqCloud APIServerless on Cloudflare network
Long contextAround 131K max1M on DeepSeek V4; 262K on Kimi

Frequently asked questions

What is the difference between Groq and Cloudflare Workers AI?

Groq trades catalog size for LPU speed and tight tail latency. Cloudflare Workers AI trades speed guarantees for 50+ open models and up to 1M context.

When should I choose Groq over Cloudflare Workers AI?

Voice agents pairing Whisper with fast replies; Strict SLAs judged on tail latency; Tight agent loops on GPT-OSS.

When should I choose Cloudflare Workers AI over Groq?

Agents that need more than 131K context; Frontier-scale open models like DeepSeek V4 Pro; LoRA adapters on smaller models.

Is Groq or Cloudflare Workers AI cheaper?

Groq: Near the floor on small models. Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. The cheaper choice depends on the model and workload.

Which has more context, Groq or Cloudflare Workers AI?

Groq: Around 131K max. Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.