vs

Z.ai vs RunInfra

Two flat-rate coding plans with five-hour and weekly quota resets. RunInfra starts cheaper on mid-size models; Z.ai's plan costs more but runs a stronger model family.

By The Subconscious Team · Updated

Z.ai vs RunInfra: key differences

The plans look alike on paper. RunInfra's coding plans start at $10 a month and Z.ai's GLM Coding Plan at $18 on the Lite tier, and both reset limits every five hours and weekly. Both plug into Claude Code, and RunInfra also lists Codex, OpenCode, Cline, Aider and dozens of other CLIs. The difference is the model. RunInfra's hosted library is tiny and centered on mid-size models like Nemotron 3.5 Lightning 30B and Qwen 3.8 27B, far from frontier quality. Z.ai serves GLM-5.3, from a family whose GLM-5 launched first among open-weight models on the Artificial Analysis index.

RunInfra has a second product that Z.ai's listing does not include: an agent that builds deployments, benchmarking a model across GPUs from L4 to B200, testing quantized variants and shipping an endpoint that scales to zero with cold starts under two seconds. It also chains models into voice pipelines. Z.ai has a lab behind it and open MIT weights, but servers mostly in China, adding 100 to 200ms from the US or Europe, and quota that burns faster during Beijing peak hours. RunInfra is young, with little independent benchmarking.

What Z.ai and RunInfra do

Z.ai

Z.AI is the international brand of Chinese lab Zhipu AI, maker of the GLM models. Its current flagship, GLM-5.3, shipped August 17, 2026 at $1.40 in and $4.40 out per million tokens, with cached input at $0.26. GLM-5.3-Flash costs $0.075 in and $0.25 out, and several older Flash models are priced at zero, a real free tier instead of trial credits. GLM-5, released in February 2026, is a 744B mixture-of-experts model under an MIT license, and at launch it ranked first among open-weight models on the Artificial Analysis index with a record-low hallucination score.

Example models: GLM-5.3, GLM-5.3-Flash

Full Z.ai profile

RunInfra

RunInfra pitches open models built for agents, with two ways in. Its hosted Model APIs serve a small curated library, including Nemotron 3.5 Lightning 30B, Qwen 3.8 27B and Ornith 1.5 35B, behind one key that works with both the OpenAI and Anthropic SDKs. Cached context bills at a discount. Coding plans start at $10 a month with limits that reset every five hours and every week, and they plug into Claude Code, Codex, OpenCode, Cline, Aider and dozens of other agent CLIs.

Example models: Nemotron 3.5 Lightning 30B, Qwen 3.8 27B

Full RunInfra profile

Should you choose Z.ai or RunInfra?

Z.ai

Choose Z.ai for

  • Stronger model quality on a flat coding plan
  • MIT weights for later self-hosting
  • Cheap per-token GLM calls outside the plan

RunInfra

Choose RunInfra for

  • The lowest-cost entry plan for agent CLIs
  • Automated deployment of a tuned open model
  • Chaining Whisper, an LLM and a TTS voice in one pipeline

Z.ai vs RunInfra at a glance

AttributeZ.aiRunInfra
Model accessOpen weights (MIT)Open weights
Flagship modelsGLM-5.3, GLM-5.3-FlashNemotron 3.5 Lightning 30B, Qwen 3.8 27B
Speed~80 tok/s on GLM-5.3Cold starts under 2s
Price$1.40 in, $4.40 out (GLM-5.3); free Flash tierCoding plans from $10 a month
CustomizationOpen weights, no license limitsUploads up to 50 GB; auto-quantization
DeploymentAPI, GLM Coding PlanModel APIs, agent-built endpoints
Long context1M (GLM-5.3)Varies by model

Frequently asked questions

What is the difference between Z.ai and RunInfra?

Two flat-rate coding plans with five-hour and weekly quota resets. RunInfra starts cheaper on mid-size models; Z.ai's plan costs more but runs a stronger model family.

When should I choose Z.ai over RunInfra?

Stronger model quality on a flat coding plan; MIT weights for later self-hosting; Cheap per-token GLM calls outside the plan.

When should I choose RunInfra over Z.ai?

The lowest-cost entry plan for agent CLIs; Automated deployment of a tuned open model; Chaining Whisper, an LLM and a TTS voice in one pipeline.

Is Z.ai or RunInfra cheaper?

Z.ai: $1.40 in, $4.40 out (GLM-5.3); free Flash tier. RunInfra: Coding plans from $10 a month. The cheaper choice depends on the model and workload.

Which has more context, Z.ai or RunInfra?

Z.ai: 1M (GLM-5.3). RunInfra: Varies by model.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.