Cloudflare Workers AI vs Z.ai
Both sell GLM 5.3 at $1.40 in and $4.40 out. Z.ai adds a flat-rate coding plan and free Flash models; Cloudflare adds servers outside China and a wider catalog.
By The Subconscious Team · Updated
Cloudflare Workers AI vs Z.ai: key differences
Price parity makes the extras decisive. Z.ai's GLM-5.3 costs $1.40 in and $4.40 out per million tokens with cached input at $0.26, and Cloudflare lists GLM 5.3 at the same rates, alongside GLM 5.2. Z.ai's cheaper tiers go further: GLM-5.3-Flash at $0.075 in and $0.25 out, and several older Flash models priced at zero. Its GLM Coding Plan starts at $18 a month on Lite with a quota that resets every five hours, which Z.ai says equals 15 to 30x the fee at API rates. Cloudflare's free allowance is 10,000 Neurons a day. Z.ai also offers an Anthropic-compatible endpoint that runs Claude Code on GLM, while Cloudflare's endpoints are OpenAI-compatible.
Location and platform tip the other way. Z.ai's servers sit mostly in China, adding 100 to 200ms from the US or Europe and raising data concerns for enterprises, and its Coding Plan quota burns 2 to 3x faster on premium models during Beijing peak hours. Cloudflare runs GLM on GPUs inside its own network, pairs it with prefix caching and session affinity for long loops, and adds 50+ other models plus AI Gateway fallbacks. GLM's MIT-licensed weights allow self-hosting or fine-tuning, which neither hosted option replaces. Individual developers on Claude Code get the most from Z.ai's plan. Teams wanting GLM in production without China-hosted data should look at Cloudflare.
What Cloudflare Workers AI and Z.ai do
Cloudflare Workers AI
Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.
Example models: DeepSeek V4 Pro, GLM 5.3
Full Cloudflare Workers AI profileZ.ai
Z.AI is the international brand of Chinese lab Zhipu AI, maker of the GLM models. Its current flagship, GLM-5.3, shipped August 17, 2026 at $1.40 in and $4.40 out per million tokens, with cached input at $0.26. GLM-5.3-Flash costs $0.075 in and $0.25 out, and several older Flash models are priced at zero, a real free tier instead of trial credits. GLM-5, released in February 2026, is a 744B mixture-of-experts model under an MIT license, and at launch it ranked first among open-weight models on the Artificial Analysis index with a record-low hallucination score.
Example models: GLM-5.3, GLM-5.3-Flash
Full Z.ai profileShould you choose Cloudflare Workers AI or Z.ai?
Cloudflare Workers AI
Choose Cloudflare Workers AI for
- GLM 5.3 without China-hosted servers
- Mixing GLM with DeepSeek and Kimi on one API
- Long GLM loops with prefix caching
Z.ai
Choose Z.ai for
- Flat-rate agentic coding on the GLM Coding Plan
- Running Claude Code on GLM
- Free Flash models for experiments
Cloudflare Workers AI vs Z.ai at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Open weights (MIT) |
| Flagship models | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B | GLM-5.3, GLM-5.3-Flash |
| Speed | Unknown | ~80 tok/s on GLM-5.3 |
| Price | $0.011 per 1K Neurons; 10K free daily | $1.40 in, $4.40 out (GLM-5.3); free Flash tier |
| Customization | BYO LoRA on small models (beta) | Open weights, no license limits |
| Deployment | Serverless on Cloudflare network | API, GLM Coding Plan |
| Long context | 1M on DeepSeek V4; 262K on Kimi | 1M (GLM-5.3) |
Frequently asked questions
What is the difference between Cloudflare Workers AI and Z.ai?
Both sell GLM 5.3 at $1.40 in and $4.40 out. Z.ai adds a flat-rate coding plan and free Flash models; Cloudflare adds servers outside China and a wider catalog.
When should I choose Cloudflare Workers AI over Z.ai?
GLM 5.3 without China-hosted servers; Mixing GLM with DeepSeek and Kimi on one API; Long GLM loops with prefix caching.
When should I choose Z.ai over Cloudflare Workers AI?
Flat-rate agentic coding on the GLM Coding Plan; Running Claude Code on GLM; Free Flash models for experiments.
Is Cloudflare Workers AI or Z.ai cheaper?
Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. Z.ai: $1.40 in, $4.40 out (GLM-5.3); free Flash tier. The cheaper choice depends on the model and workload.
Which has more context, Cloudflare Workers AI or Z.ai?
Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi. Z.ai: 1M (GLM-5.3).
Related comparisons
Subconscious vs Cloudflare Workers AI
OpenAI vs Cloudflare Workers AI
Anthropic vs Cloudflare Workers AI
Google Vertex AI vs Cloudflare Workers AI
Amazon Bedrock vs Cloudflare Workers AI
Together AI vs Cloudflare Workers AI
Subconscious vs Z.ai
OpenAI vs Z.ai
Anthropic vs Z.ai
Google Vertex AI vs Z.ai
Amazon Bedrock vs Z.ai
Together AI vs Z.ai
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.