Cloudflare Workers AI vs Inference.net
Inference.net runs cheap batch on spare GPU capacity and distills traffic into custom models. Workers AI serves real-time open models from Cloudflare's network.
By The Subconscious Team · Updated
Cloudflare Workers AI vs Inference.net: key differences
Latency tolerance is the first filter. Inference.net's Batch API takes up to 1M requests per file with completion windows from 24 hours to 7 days, priced off discounted spare GPU time, and the company says fragmented capacity fits batch better than strict real-time SLAs. Workers AI is synchronous serverless inference, called from a Worker or over REST, with per-token equivalents like gpt-oss 120B at $0.35 in and $0.75 out and 10,000 free Neurons a day. Inference.net publishes few public price comparisons, so buyers lean on its numbers.
The second filter is how much you want to own the model. Inference.net's gateway captures production traffic, turns it into eval and training sets, fine-tunes a task-specific model and hosts it on a dedicated GPU with a 99.99% uptime target. Workers AI offers no fine-tuning or dedicated hosting for large LLMs, only a beta small-model LoRA. What Cloudflare gives instead is a ready catalog of 50+ models, including Kimi K2.7 Code and DeepSeek V4 with 1M context, inside the same platform as Workers, storage and AI Gateway.
What Cloudflare Workers AI and Inference.net do
Cloudflare Workers AI
Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.
Example models: DeepSeek V4 Pro, GLM 5.3
Full Cloudflare Workers AI profileInference.net
Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.
Example models: open catalog models plus customer fine-tunes served on dedicated GPUs
Full Inference.net profileShould you choose Cloudflare Workers AI or Inference.net?
Cloudflare Workers AI
Choose Cloudflare Workers AI for
- Real-time chat and agents on Workers
- Frontier open models with no training step
- A free tier for early prototypes
Inference.net
Choose Inference.net for
- Offline extraction and synthetic data at scale
- Distilling a closed-API workload into a small model
- Dedicated GPUs for a custom fine-tune
Cloudflare Workers AI vs Inference.net at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Open, closed and custom |
| Flagship models | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B | Customer fine-tunes |
| Speed | Unknown | Batch windows of 24h to 7 days |
| Price | $0.011 per 1K Neurons; 10K free daily | Discounted spare GPU capacity |
| Customization | BYO LoRA on small models (beta) | Distill traces into custom models |
| Deployment | Serverless on Cloudflare network | Batch API, gateway, dedicated GPUs |
| Long context | 1M on DeepSeek V4; 262K on Kimi | Varies by model |
Frequently asked questions
What is the difference between Cloudflare Workers AI and Inference.net?
Inference.net runs cheap batch on spare GPU capacity and distills traffic into custom models. Workers AI serves real-time open models from Cloudflare's network.
When should I choose Cloudflare Workers AI over Inference.net?
Real-time chat and agents on Workers; Frontier open models with no training step; A free tier for early prototypes.
When should I choose Inference.net over Cloudflare Workers AI?
Offline extraction and synthetic data at scale; Distilling a closed-API workload into a small model; Dedicated GPUs for a custom fine-tune.
Is Cloudflare Workers AI or Inference.net cheaper?
Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. Inference.net: Discounted spare GPU capacity. The cheaper choice depends on the model and workload.
Which has more context, Cloudflare Workers AI or Inference.net?
Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi. Inference.net: Varies by model.
Related comparisons
Subconscious vs Cloudflare Workers AI
OpenAI vs Cloudflare Workers AI
Anthropic vs Cloudflare Workers AI
Google Vertex AI vs Cloudflare Workers AI
Amazon Bedrock vs Cloudflare Workers AI
Together AI vs Cloudflare Workers AI
Subconscious vs Inference.net
OpenAI vs Inference.net
Anthropic vs Inference.net
Google Vertex AI vs Inference.net
Amazon Bedrock vs Inference.net
Together AI vs Inference.net
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.