We raised $5.1M for long-running agents.
vs

Z.ai vs Cohere

Z.ai's GLM models are cheap, MIT-licensed and strong at coding. Cohere's Command models cost more but come with retrieval tools and private deployment support.

By The Subconscious Team · Updated

Z.ai vs Cohere: key differences

Z.ai competes on price and coding. GLM-5.3 costs $1.40 in and $4.40 out with a 1M window, GLM-5.3-Flash is $0.075 in and $0.25 out, and several older Flash models are free. The $18-a-month GLM Coding Plan and an Anthropic-compatible endpoint made GLM a popular cheap backend for Claude Code. Cohere's Command A lists at $2.50 in and $10 out with 256K context, and Cohere acknowledges Command A+ trails the latest GLM models on agentic coding. Both publish permissive weights, MIT for GLM-5 and Apache 2.0 for Command A+.

Enterprise posture divides them. Z.ai's servers sit mostly in China, adding 100 to 200ms from the US or Europe and raising data concerns, and Coding Plan quota burns faster during Beijing peak hours. Cohere sells to banks and governments, runs on Bedrock, Azure AI Foundry and OCI, and supports VPC and on-prem deployment with fine-tuning. Embed 4 and Rerank 4 give Cohere a retrieval stack Z.ai lacks. Because GLM weights are open, teams can also get GLM from US hosts. For budget coding agents, Z.ai. For governed enterprise search, Cohere.

What Z.ai and Cohere do

Z.ai

Z.AI is the international brand of Chinese lab Zhipu AI, maker of the GLM models. Its current flagship, GLM-5.3, shipped August 17, 2026 at $1.40 in and $4.40 out per million tokens, with cached input at $0.26. GLM-5.3-Flash costs $0.075 in and $0.25 out, and several older Flash models are priced at zero, a real free tier instead of trial credits. GLM-5, released in February 2026, is a 744B mixture-of-experts model under an MIT license, and at launch it ranked first among open-weight models on the Artificial Analysis index with a record-low hallucination score.

Example models: GLM-5.3, GLM-5.3-Flash

Full Z.ai profile

Cohere

Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.

Example models: Command A+, Command A, Embed 4, Rerank 4

Full Cohere profile

Should you choose Z.ai or Cohere?

Z.ai

Choose Z.ai for

  • Budget coding inside Claude Code
  • 1M context at low per-token prices
  • A real free tier on Flash models

Cohere

Choose Cohere for

  • Western enterprise procurement and support
  • On-prem RAG with fine-tuning
  • Retrieval models alongside generation

Z.ai vs Cohere at a glance

AttributeZ.aiCohere
Model accessOpen weights (MIT)Closed, plus open Command A+
Flagship modelsGLM-5.3, GLM-5.3-FlashCommand A+, Command A, Embed 4, Rerank 4
Speed~80 tok/s on GLM-5.3375 tok/s on Command A+ W4A4, per Cohere
Price$1.40 in, $4.40 out (GLM-5.3); free Flash tier$0.0375–$2.50 in, $0.15–$10 out per 1M
CustomizationOpen weights, no license limitsEnterprise fine-tuning, incl. private
DeploymentAPI, GLM Coding PlanAPI, Bedrock, Azure, OCI, VPC, on-prem
Long context1M (GLM-5.3)256K on Command A; 128K on A+

Frequently asked questions

What is the difference between Z.ai and Cohere?

Z.ai's GLM models are cheap, MIT-licensed and strong at coding. Cohere's Command models cost more but come with retrieval tools and private deployment support.

When should I choose Z.ai over Cohere?

Budget coding inside Claude Code; 1M context at low per-token prices; A real free tier on Flash models.

When should I choose Cohere over Z.ai?

Western enterprise procurement and support; On-prem RAG with fine-tuning; Retrieval models alongside generation.

Is Z.ai or Cohere cheaper?

Z.ai: $1.40 in, $4.40 out (GLM-5.3); free Flash tier. Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. The cheaper choice depends on the model and workload.

Which has more context, Z.ai or Cohere?

Z.ai: 1M (GLM-5.3). Cohere: 256K on Command A; 128K on A+.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.