Cloudflare Workers AI vs Luminal
Workers AI serves a fixed open-model catalog on Cloudflare GPUs, called from Workers. Luminal compiles models you bring into faster GPU code.
By The Subconscious Team · Updated
Cloudflare Workers AI vs Luminal: key differences
Workers AI runs a set catalog, including DeepSeek V4 Pro, GLM 5.3 and gpt-oss 120B, on Cloudflare's network, priced in Neurons with 10K free daily, and callable straight from Workers. Custom work is limited to LoRA on small models. Luminal goes the other way: you bring the model, and its compiler emits fused native kernels ahead of time for serverless or on-prem serving.
Workers AI is the easy pick for apps already on Cloudflare that want a free tier and edge integration. Luminal is for teams with a specific model, possibly a custom one, that want maximum throughput; it reports 36K tokens per second on GPT-OSS 120B across 8 H100s. Its cloud is early access and unpriced.
What Cloudflare Workers AI and Luminal do
Cloudflare Workers AI
Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.
Example models: DeepSeek V4 Pro, GLM 5.3
Full Cloudflare Workers AI profileLuminal
Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.
Example models: GPT-OSS 120B, Llama 3 8B
Full Luminal profileShould you choose Cloudflare Workers AI or Luminal?
Cloudflare Workers AI
Choose Cloudflare Workers AI for
- Inference called directly from Workers
- A free daily allowance
- No infrastructure to manage
Luminal
Choose Luminal for
- Serving custom or fine-tuned architectures off any catalog
- Maximum throughput per GPU on a self-chosen model
- On-prem deployments with custom kernel work and SLAs
Cloudflare Workers AI vs Luminal at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Bring your own weights |
| Flagship models | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B | No public catalog |
| Speed | Unknown | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | $0.011 per 1K Neurons; 10K free daily | Pay per use; rates not published |
| Customization | BYO LoRA on small models (beta) | Compiles any PyTorch or HF model |
| Deployment | Serverless on Cloudflare network | Serverless (early access), on-prem license |
| Long context | 1M on DeepSeek V4; 262K on Kimi | Unknown |
Frequently asked questions
What is the difference between Cloudflare Workers AI and Luminal?
Workers AI serves a fixed open-model catalog on Cloudflare GPUs, called from Workers. Luminal compiles models you bring into faster GPU code.
When should I choose Cloudflare Workers AI over Luminal?
Inference called directly from Workers; A free daily allowance; No infrastructure to manage.
When should I choose Luminal over Cloudflare Workers AI?
Serving custom or fine-tuned architectures off any catalog; Maximum throughput per GPU on a self-chosen model; On-prem deployments with custom kernel work and SLAs.
Is Cloudflare Workers AI or Luminal cheaper?
Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.
Related comparisons
Subconscious vs Cloudflare Workers AI
OpenAI vs Cloudflare Workers AI
Anthropic vs Cloudflare Workers AI
Google Vertex AI vs Cloudflare Workers AI
Amazon Bedrock vs Cloudflare Workers AI
Together AI vs Cloudflare Workers AI
Subconscious vs Luminal
OpenAI vs Luminal
Anthropic vs Luminal
Google Vertex AI vs Luminal
Amazon Bedrock vs Luminal
Together AI vs Luminal
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.