Cloudflare Workers AI
Serverless open-model inference on Cloudflare GPUs, called straight from Workers and billed in Neurons.
- Founded
- 2009
- Example models
- DeepSeek V4 Pro, GLM 5.3
By The Subconscious Team · Updated
What is Cloudflare Workers AI?
Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.
Billing is in Neurons at $0.011 per 1,000, with 10,000 Neurons a day free, and Cloudflare publishes per-token equivalents for each model. DeepSeek V4 Pro costs $1.32 in and $3.96 out per million, GLM 5.3 $1.40 and $4.40, Kimi K2.6 $0.95 and $4.00, and gpt-oss 120B $0.35 and $0.75. Prefix caching bills cached input at a discount, and an x-session-affinity header pins a conversation to one instance to raise hit rates. LoRA adapters are bring-your-own, in free open beta, capped at rank 32, 300MB and 100 per account, and limited to smaller non-quantized models. AI Gateway adds caching, rate limits, retries, fallbacks and logs, and since August 2026 shares billing with Workers AI.
Cloudflare Workers AI pros, cons and use cases
Upsides
- Inference, code, storage and the Agents SDK live on one platform, so an agent can run end to end on Cloudflare.
- Daily free allocation and per-token prices that undercut many GPU hosts on gpt-oss and Llama models.
- OpenAI-compatible endpoints make it a base URL swap for existing SDK code.
- AI Gateway in front of Workers AI and third-party providers gives one place for caching, fallbacks and spend tracking.
Core use cases
- Agents and chat features built on Workers that need an LLM without a separate vendor.
- Low-cost classification, embeddings and summarization inside existing Cloudflare apps.
- Long-context agent loops on DeepSeek V4 or GLM 5.3 with prefix caching.
Downsides
- LoRA covers only smaller, older models, and there is no fine-tuning or dedicated deployment for the large LLMs.
- Large models like Kimi K2.6, K2.7 Code and GLM 5.2 require the Workers Paid plan, and synchronous requests can queue for capacity.
Cloudflare Workers AI alternatives compared
Pick any row for the full head-to-head.
| Provider | Model access | Flagship models | Speed | Price | Customization | Deployment | Long context | Compare |
|---|---|---|---|---|---|---|---|---|
| Open weights | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B | Unknown | $0.011 per 1K Neurons; 10K free daily | BYO LoRA on small models (beta) | Serverless on Cloudflare network | 1M on DeepSeek V4; 262K on Kimi | ||
| Open weights | GLM 5.3, DeepSeek V4.1 Flash | 2x faster task completion | 50–80% lower cost; billed on processed tokens | Marathon post-trained variants | Managed API, dedicated, on-prem | 5M+ effective context | Compare | |
| Closed, plus open gpt-oss | GPT-6 Astra, GPT-5.6 Sol, Terra, Luna | Fast mode: up to 2.5x at 2x price | $0.20–$10 in, $1.20–$50 out per 1M | N/A | API, Azure OpenAI, Bedrock | 1.05M; 2x input past 272K | Compare | |
| Closed | Claude Fable 5.1, Opus, Sonnet, Haiku 4.5 | Fable is the slowest tier | $1–$10 in, $5–$50 out per 1M | N/A | API, Bedrock, Vertex AI, Microsoft Foundry | 1M, no surcharge past 200K | Compare | |
| Closed and open, 200+ models | Gemini 3.8 Flash, Claude, Gemma | Flash tier built for low latency | Gemini 3.8 Flash $0.75 in, $3.75 out | Custom training on GPUs or TPUs | Managed on Google Cloud | 1M on Gemini 3.8 Flash | Compare | |
| Closed and open, 100+ models | Claude, GPT-6 Astra, Nova, DeepSeek | Latency-optimized option on some models | ~20–35% above direct; Claude at parity | Fine-tuning, Custom Model Import | Managed on AWS, AgentCore | Varies by model | Compare | |
| Open weights | Kimi K3, DeepSeek V4, GLM 5.2, Qwen 3.8 | 0.99s TTFT on DeepSeek V4 Pro | Parity with Fireworks and Baseten | LoRA and full SFT; RL in beta | Serverless, dedicated, GPU clusters | 512K on DeepSeek V4 Pro | Compare | |
| Open weights | DeepSeek V4 Pro, Kimi K3 | 167–174 tok/s on DeepSeek V4 Pro | Fine-tunes served at base price | SFT, DPO, RFT; Training API | Serverless, dedicated GPUs | Full 1M on DeepSeek V4 Pro | Compare | |
| Open weights, 13 curated | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | 0.49s TTFT, lowest measured | H100 about $6.50/hr dedicated | Deploy any model with Truss | Model APIs, dedicated, self-host | Varies by model | Compare | |
| Open weights | GPT-OSS 120B, Qwen 3.6 27B | 500–1,000 tok/s | Near the floor on small models | No fine-tuned model hosting | GroqCloud API | Around 131K max | Compare | |
| Open weights | GPT-OSS 120B, Gemma 4 31B | ~3,000 tok/s on GPT-OSS 120B | $0.35 in, $0.75 out (GPT-OSS 120B) | Unknown | Shared API, dedicated, partners | Unknown | Compare | |
| Open weights | DeepSeek V4 Flash, Llama 3.1 8B | ~33 tok/s on DeepSeek V4 Pro (FP4) | From $0.02 per 1M | No managed fine-tuning | Shared API, no contracts | 66K on FP4 DeepSeek V4 Pro | Compare | |
| Open weights | GLM 5.3, Kimi K3, DeepSeek V4.1 Flash | Routes to fastest provider by default | Provider rates, no markup | N/A | Serverless router; dedicated Endpoints | Up to 1M, provider-dependent | Compare | |
| Bring your own weights | None hosted | ~1s container boot | Per second; H100 $3.95/hr list | Run any training code | Serverless GPU containers | Depends on the model you deploy | Compare | |
| Closed | Grok 4.6, Grok 4.20, grok-build | ~54 tok/s on Grok 4.6 | $2 in, $6 out (Grok 4.6); 2x past 200K | Unknown | First-party API | 500K (4.6), 1M (4.20, 4.3) | Compare | |
| Open weights, plus closed Codestral | Mistral Medium 3.5, Small 4, Large 3 | Unknown | $0.15–$1.50 in, $0.60–$7.50 out per 1M | Forge (enterprise); fine-tuning API deprecated | API, Azure, Bedrock, Vertex, self-host | 256K | Compare | |
| Open weights (MIT) | DeepSeek V4.1 Flash, V4 Pro | ~35 tok/s on V4 Pro | Off-peak hours at half price | Open weights to fine-tune | First-party API, Hugging Face weights | 1M, 384K max output | Compare | |
| Open weights, custom license | Kimi K3, Kimi K2.6 | ~33 tok/s on Kimi K3 | $3 in, $15 out (Kimi K3) | Open weights to fine-tune | API, Kimi Code, OpenRouter | 1M | Compare | |
| Open weights (MIT) | GLM-5.3, GLM-5.3-Flash | ~80 tok/s on GLM-5.3 | $1.40 in, $4.40 out (GLM-5.3); free Flash tier | Open weights, no license limits | API, GLM Coding Plan | 1M (GLM-5.3) | Compare | |
| Closed Max; open smaller Qwen | Qwen 3.8-Max, Qwen 3.7-Max | ~40 tok/s on Qwen 3.8-Max | $2 in, $6 out international | No fine-tuning on Max | Model Studio on Alibaba Cloud | 1M (Qwen 3.8-Max) | Compare | |
| Closed API; open Muse Glimmer | Muse Spark 1.3, Muse Glimmer | ~145–233 tok/s on Muse Spark 1.3 | $1.25 in, $4.25 out; Contributor tier cheaper | Open Muse Glimmer weights to fine-tune | Meta Model API (preview) | 1M | Compare | |
| Closed, plus open Command A+ | Command A+, Command A, Embed 4, Rerank 4 | 375 tok/s on Command A+ W4A4, per Cohere | $0.0375–$2.50 in, $0.15–$10 out per 1M | Enterprise fine-tuning, incl. private | API, Bedrock, Azure, OCI, VPC, on-prem | 256K on Command A; 128K on A+ | Compare | |
| Open weights | MiniMax M2.7, GPT-OSS 120B, DeepSeek | ~820 tok/s on MiniMax M2.7 (SN50) | $0.22 in, $0.59 out (GPT-OSS 120B) | Unknown | SambaCloud, racks for neoclouds | Up to 192K (MiniMax M2.7) | Compare | |
| Open weights, 60+ models | DeepSeek, Qwen, GLM, Kimi, GPT-OSS | Among top hosts on throughput | From $0.06 per 1M input | Serve uploaded fine-tunes | Token Factory, dedicated, raw GPUs | Varies by model | Compare | |
| Open weights | DeepSeek V4, GLM 5.3, Kimi K2.6, Nemotron 3 | Up to 9.9x faster TTFT vs vLLM (vendor claim) | $0.05–$1.74 in, $0.20–$4.40 out per 1M | Serverless LoRA fine-tuning | Serverless, self-serve and tailored dedicated, raw GPUs | Varies by model; cluster-wide KV cache | Compare | |
| Hosted media models | FLUX, Kling, Seedream | Cold starts on less popular endpoints | Per image, per video second, GPU time | LoRA training endpoints | Hosted API, serverless GPUs | Not applicable | Compare | |
| Open weights | DeepSeek V4 Pro, Gemma 4 | ~36 tok/s on DeepSeek V4 Pro | From $0.02 per 1M; batch 50% off | Hot-swappable LoRA adapters | Serverless, GPU cloud, dedicated | Full 1M on DeepSeek V4 Pro | Compare | |
| Open weights, plus proxied closed models | GLM 5.3, Kimi K3, DeepSeek V4 Pro | Unknown | $0.06–$12 in, $0.28–$60 out per 1M; DIEM staking | Unknown | Serverless API, consumer app | 1M on most current models | Compare | |
| Any Hugging Face model | GTE-Qwen2, Qwen3-VL-8B-Instruct | 600ms p99 real-time budget | Per-parameter rates; batch 50% off | Private Hugging Face repos | Serverless, elastic, dedicated, batch | Varies by model | Compare | |
| Open, closed and custom | Customer fine-tunes | Batch windows of 24h to 7 days | Discounted spare GPU capacity | Distill traces into custom models | Batch API, gateway, dedicated GPUs | Varies by model | Compare | |
| Open and third-party models | GLM-4.7-Flash, Google Veo | Near bare-metal performance | $0.07 in, $0.40 out (GLM-4.7-Flash) | Unknown | Shared, autoscaling, reserved GPUs | Varies by model | Compare | |
| Thinking Machines | Open weights | Inkling, Inkling-Small | Unknown | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | LoRA SFT and RL via Tinker | Training API, beta serverless (Inkling only) | Inkling up to 1M; Tinker 32K–256K | Compare |
| Open weights | Kimi K2.6, GLM-5, GPT-OSS 120B | Minutes per turn by design | 30–80% off by completion window | Customer LoRA fine-tunes | API plus Sailboxes | Varies by model | Compare | |
| Specialist models | morph-v3-fast, morph-v3-large | 10,500+ tok/s Fast Apply | ~40% fewer tokens than full rewrites | Fine-tuning offered | OpenAI-compatible API | Unknown | Compare | |
| Specialist models | relace-apply-3, agentic search | ~10,000 tok/s apply | 3x+ cheaper than full rewrites | Unknown | Hosted API or self-hosted | 128K max | Compare | |
| Decision models | Jev, jev-1.13 | ~100ms per call | A fraction of an LLM call | Unknown | Early-access API | Unknown | Compare | |
| Open (Apache 2.0) and API models | Step 3.7 Flash, Step3 | ~128 tok/s on Step 3.7 Flash | $0.20 in, $1.15 out (Step 3.7 Flash) | Open weights to fine-tune | First-party API, OpenRouter | 256K | Compare | |
| Hosted media models | Seedance 2.5, Qwen-Image-3.0 | Unknown | Images from fractions of a cent | Fine-tuned diffusion checkpoints | Unified API, raw GPUs | Not applicable | Compare | |
| Proprietary coding models | KAT-Coder-Pro V2.5, KAT-Coder-Air | Unknown | Per token or KwaiKAT Coding Plan | Unknown | MaaS API, bare metal | Unknown | Compare | |
| Open weights | Qwen 3.5 397B Turbo, GLM 5.1 Turbo | 2–2.8x vs stock vLLM or SGLang | Wafer Pass from $10 a week | Agent-tuned dedicated deployments | Serverless pass, dedicated | Varies by model | Compare | |
| Open weights | Nemotron 3.5 Lightning 30B, Qwen 3.8 27B | Cold starts under 2s | Coding plans from $10 a month | Uploads up to 50 GB; auto-quantization | Model APIs, agent-built endpoints | Varies by model | Compare | |
| Open weights | DeepSeek V4.1 Flash, GLM 5.3 Flash | ~157 tok/s on DeepSeek V4.1 Flash | $0.10 in, $0.40 out (GLM 5.3 Flash) | Unknown | Via Vercel AI Gateway | 1M | Compare |
Frequently asked questions
What is Cloudflare Workers AI?
Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.
What is Cloudflare Workers AI best for?
Agents and chat features built on Workers that need an LLM without a separate vendor; Low-cost classification, embeddings and summarization inside existing Cloudflare apps; Long-context agent loops on DeepSeek V4 or GLM 5.3 with prefix caching.
How much does Cloudflare Workers AI cost?
Cloudflare Workers AI pricing at a glance: $0.011 per 1K Neurons; 10K free daily. Rates change often, so check Cloudflare Workers AI's pricing page before committing.
How much context does Cloudflare Workers AI support?
Cloudflare Workers AI's long-context support: 1M on DeepSeek V4; 262K on Kimi.
What are the downsides of Cloudflare Workers AI?
LoRA covers only smaller, older models, and there is no fine-tuning or dedicated deployment for the large LLMs; Large models like Kimi K2.6, K2.7 Code and GLM 5.2 require the Workers Paid plan, and synchronous requests can queue for capacity.
What are the best alternatives to Cloudflare Workers AI?
Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Cloudflare Workers AI on this site.
Sources: Workers AI pricing, Workers AI changelog, Workers AI now runs large models, Workers AI LoRA adapters. Pricing and model lineups change often; figures are a snapshot.