vs

Z.ai vs StreamLake

Two Chinese coding plans that drop into Claude Code. Z.ai's GLM is MIT-licensed with dollar pricing; StreamLake's KAT-Coder is proprietary and China-first.

By The Subconscious Team · Updated

Z.ai vs StreamLake: key differences

Both companies court the same developer: someone who wants Claude Code-style agentic coding for less. Z.ai offers the GLM Coding Plan from $18 a month on the Lite tier, with quota that resets every five hours and weekly, and an Anthropic-compatible endpoint. StreamLake offers the KwaiKAT Coding Plan or per-token billing on KAT-Coder-Pro V2.5, with OpenAI-protocol endpoints and a Claude-protocol proxy for Claude Code or OpenClaw. The models differ in openness. GLM weights are MIT-licensed and run on most major hosts, while KAT-Coder is proprietary to Kuaishou.

Both carry China-related trade-offs, to different degrees. Z.ai's servers sit mostly in China, adding 100 to 200ms from the US or Europe, and its quota burns 2 to 3x faster on premium models during Beijing peak hours. StreamLake's pricing and much of its documentation lead with China and yuan, and data residency in China rules it out for many US and EU buyers. Z.ai publishes dollar prices and has a track record behind GLM-5, which launched first among open-weight models on the Artificial Analysis index. Open weights also let GLM users switch hosts later.

What Z.ai and StreamLake do

Z.ai

Z.AI is the international brand of Chinese lab Zhipu AI, maker of the GLM models. Its current flagship, GLM-5.3, shipped August 17, 2026 at $1.40 in and $4.40 out per million tokens, with cached input at $0.26. GLM-5.3-Flash costs $0.075 in and $0.25 out, and several older Flash models are priced at zero, a real free tier instead of trial credits. GLM-5, released in February 2026, is a 744B mixture-of-experts model under an MIT license, and at launch it ranked first among open-weight models on the Artificial Analysis index with a record-low hallucination score.

Example models: GLM-5.3, GLM-5.3-Flash

Full Z.ai profile

StreamLake

StreamLake is the AI cloud brand of Kuaishou, the Chinese short-video company behind the Kling video models. It sells model-as-a-service inference and bare-metal compute to internet businesses, drawing on the infrastructure Kuaishou built to serve video at massive scale. Its developer site offers APIs, SDKs and integration guides aimed at taking teams from testing to production.

Example models: KAT-Coder-Pro V2.5, KAT-Coder-Air

Full StreamLake profile

Should you choose Z.ai or StreamLake?

Z.ai

Choose Z.ai for

  • Western developers who want dollar pricing and a flat plan
  • Teams that may move to self-hosting later
  • Budget coding in Claude Code

StreamLake

Choose StreamLake for

  • Developers who want Kuaishou's KAT-Coder specifically
  • Domestic Chinese MaaS plus bare-metal servers
  • Teams comfortable with yuan-first procurement

Z.ai vs StreamLake at a glance

AttributeZ.aiStreamLake
Model accessOpen weights (MIT)Proprietary coding models
Flagship modelsGLM-5.3, GLM-5.3-FlashKAT-Coder-Pro V2.5, KAT-Coder-Air
Speed~80 tok/s on GLM-5.3Unknown
Price$1.40 in, $4.40 out (GLM-5.3); free Flash tierPer token or KwaiKAT Coding Plan
CustomizationOpen weights, no license limitsUnknown
DeploymentAPI, GLM Coding PlanMaaS API, bare metal
Long context1M (GLM-5.3)Unknown

Frequently asked questions

What is the difference between Z.ai and StreamLake?

Two Chinese coding plans that drop into Claude Code. Z.ai's GLM is MIT-licensed with dollar pricing; StreamLake's KAT-Coder is proprietary and China-first.

When should I choose Z.ai over StreamLake?

Western developers who want dollar pricing and a flat plan; Teams that may move to self-hosting later; Budget coding in Claude Code.

When should I choose StreamLake over Z.ai?

Developers who want Kuaishou's KAT-Coder specifically; Domestic Chinese MaaS plus bare-metal servers; Teams comfortable with yuan-first procurement.

Is Z.ai or StreamLake cheaper?

Z.ai: $1.40 in, $4.40 out (GLM-5.3); free Flash tier. StreamLake: Per token or KwaiKAT Coding Plan. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.