Long-running agents deserve better inference.
vs

Z.ai vs Luminal

Z.ai makes MIT-licensed GLM models with a cheap coding plan. Luminal compiles open models into native GPU code for self-hosting.

By The Subconscious Team · Updated

Z.ai vs Luminal: key differences

Z.ai serves GLM-5.3 and a free Flash tier on its own API and sells a flat-rate GLM Coding Plan, with weights under MIT. Servers sit mostly in China, which adds latency from the US and Europe. Luminal sells an engine: its compiler takes a model you bring and emits native GPU kernels ahead of time, served serverless in early access or on-prem.

Teams happy with China-hosted inference get GLM cheapest from Z.ai. Teams that want GLM weights running in their own region could look at Luminal as the serving layer; it reports 36K tokens per second on GPT-OSS 120B across 8 H100s but has not published GLM figures or prices.

What Z.ai and Luminal do

Z.ai

Z.AI is the international brand of Chinese lab Zhipu AI, maker of the GLM models. Its current flagship, GLM-5.3, shipped August 17, 2026 at $1.40 in and $4.40 out per million tokens, with cached input at $0.26. GLM-5.3-Flash costs $0.075 in and $0.25 out, and several older Flash models are priced at zero, a real free tier instead of trial credits. GLM-5, released in February 2026, is a 744B mixture-of-experts model under an MIT license, and at launch it ranked first among open-weight models on the Artificial Analysis index with a record-low hallucination score.

Example models: GLM-5.3, GLM-5.3-Flash

Full Z.ai profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Z.ai or Luminal?

Z.ai

Choose Z.ai for

  • Cheap flat-rate GLM coding plan
  • A free Flash tier
  • MIT weights with no license limits

Luminal

Choose Luminal for

  • Self-hosting open weights in your own region
  • Maximum throughput per GPU on a self-chosen model
  • An open-source engine teams can run on their own hardware

Z.ai vs Luminal at a glance

AttributeZ.aiLuminal
Model accessOpen weights (MIT)Bring your own weights
Flagship modelsGLM-5.3, GLM-5.3-FlashNo public catalog
Speed~80 tok/s on GLM-5.336K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price$1.40 in, $4.40 out (GLM-5.3); free Flash tierPay per use; rates not published
CustomizationOpen weights, no license limitsCompiles any PyTorch or HF model
DeploymentAPI, GLM Coding PlanServerless (early access), on-prem license
Long context1M (GLM-5.3)Unknown

Frequently asked questions

What is the difference between Z.ai and Luminal?

Z.ai makes MIT-licensed GLM models with a cheap coding plan. Luminal compiles open models into native GPU code for self-hosting.

When should I choose Z.ai over Luminal?

Cheap flat-rate GLM coding plan; A free Flash tier; MIT weights with no license limits.

When should I choose Luminal over Z.ai?

Self-hosting open weights in your own region; Maximum throughput per GPU on a self-chosen model; An open-source engine teams can run on their own hardware.

Is Z.ai or Luminal cheaper?

Z.ai: $1.40 in, $4.40 out (GLM-5.3); free Flash tier. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.