# Z.ai vs Inference.net

> Z.ai supplies cheap GLM models; Inference.net supplies batch on spare GPUs and a pipeline from traces to custom models. They cut cost at different layers.

Canonical: https://www.subconscious.dev/compare/z-ai-vs-inference-net · By The Subconscious Team · Updated September 30, 2026

## How they compare

Both attack inference cost, at different layers. Z.ai lowers it at the model level: GLM-5.3 at $1.40 in and $4.40 out, GLM-5.3-Flash at $0.075 in and $0.25 out, and older Flash models at zero. Inference.net lowers it at the infrastructure and workflow level. Its OpenAI-compatible Batch API runs on spare GPU capacity bought at steep discounts, taking up to 1M requests per file with completion windows from 24 hours to 7 days. Its Inference Gateway routes traffic to open, closed or custom models and captures it as eval and training data.

The two can work in sequence. A team might start a narrow task on GLM, capture that traffic through a gateway, then have Inference.net fine-tune a smaller task-specific model and deploy it on a dedicated GPU with a 99.99% uptime target. GLM's MIT license puts no limits on that kind of downstream use. Inference.net's weak point is evidence, with few independent benchmarks or public pricing comparisons. Z.ai's weak points are servers mostly in China and 100 to 200ms of added latency from the US or Europe.

## What each one does

### Z.ai

Z.AI is the international brand of Chinese lab Zhipu AI, maker of the GLM models. Its current flagship, GLM-5.3, shipped August 17, 2026 at $1.40 in and $4.40 out per million tokens, with cached input at $0.26. GLM-5.3-Flash costs $0.075 in and $0.25 out, and several older Flash models are priced at zero, a real free tier instead of trial credits. GLM-5, released in February 2026, is a 744B mixture-of-experts model under an MIT license, and at launch it ranked first among open-weight models on the Artificial Analysis index with a record-low hallucination score.

### Inference.net

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

## Which is best, and when

### Choose Z.ai for

- Cheap real-time GLM calls and coding
- Free-tier experiments
- MIT weights to distill or fine-tune freely

### Choose Inference.net for

- Large offline jobs on discounted spare capacity
- Turning production traffic into a custom model
- Batch windows of 24 hours to 7 days

## At a glance

| Attribute | Z.ai | Inference.net |
|---|---|---|
| Model access | Open weights (MIT) | Open, closed and custom |
| Flagship models | GLM-5.3, GLM-5.3-Flash | Customer fine-tunes |
| Speed | ~80 tok/s on GLM-5.3 | Batch windows of 24h to 7 days |
| Price | $1.40 in, $4.40 out (GLM-5.3); free Flash tier | Discounted spare GPU capacity |
| Customization | Open weights, no license limits | Distill traces into custom models |
| Deployment | API, GLM Coding Plan | Batch API, gateway, dedicated GPUs |
| Long context | 1M (GLM-5.3) | Varies by model |

## FAQ

### What is the difference between Z.ai and Inference.net?

Z.ai supplies cheap GLM models; Inference.net supplies batch on spare GPUs and a pipeline from traces to custom models. They cut cost at different layers.

### When should I choose Z.ai over Inference.net?

Cheap real-time GLM calls and coding; Free-tier experiments; MIT weights to distill or fine-tune freely.

### When should I choose Inference.net over Z.ai?

Large offline jobs on discounted spare capacity; Turning production traffic into a custom model; Batch windows of 24 hours to 7 days.

### Is Z.ai or Inference.net cheaper?

Z.ai: $1.40 in, $4.40 out (GLM-5.3); free Flash tier. Inference.net: Discounted spare GPU capacity. The cheaper choice depends on the model and workload.

### Which has more context, Z.ai or Inference.net?

Z.ai: 1M (GLM-5.3). Inference.net: Varies by model.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Z.ai](https://www.subconscious.dev/compare/subconscious-vs-z-ai.md), [Subconscious vs Inference.net](https://www.subconscious.dev/compare/subconscious-vs-inference-net.md).

Full profiles: [Z.ai](https://www.subconscious.dev/providers/z-ai.md), [Inference.net](https://www.subconscious.dev/providers/inference-net.md).
