# Z.ai vs Thinking Machines

> Z.ai sells GLM-5.3 cheaply with a flat-rate coding plan. Thinking Machines lets teams post-train GLM-5.3 with Tinker and offers its own Inkling models.

Canonical: https://www.subconscious.dev/compare/z-ai-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Z.ai makes GLM, and GLM-5.3 is one of the bases Thinking Machines' Tinker can train. Z.ai's API charges $1.40 in and $4.40 out on GLM-5.3 with 1M context, GLM-5.3-Flash costs $0.075 in and $0.25 out, and several older Flash models are free. The GLM Coding Plan starts at $18 a month and works inside Claude Code through an Anthropic-compatible endpoint. Weights are MIT-licensed. Thinking Machines does not compete on cheap inference. Tinker lets teams write SFT or RL loops with LoRA adapters, billed per million prefill, sample and train tokens, and its Inkling models sell at $1.00 in and $4.05 out on a beta serverless API.

Z.ai wins on price and developer tooling. The free Flash tier and flat coding plan are hard to beat for budget agentic coding. Its downsides are servers mostly in China, adding 100 to 200ms from the US or Europe and raising data concerns, plus quota that burns faster during Beijing peak hours. Thinking Machines is a US lab with Apache 2.0 Inkling models that add image and audio input and up to 1M context, but it only serves Inkling, and checkpoint sampling is for testing. GLM's MIT weights can be fine-tuned anywhere; Tinker is the option for teams that want RL on GLM-5.3 without running GPUs.

## What each one does

### Z.ai

Z.AI is the international brand of Chinese lab Zhipu AI, maker of the GLM models. Its current flagship, GLM-5.3, shipped August 17, 2026 at $1.40 in and $4.40 out per million tokens, with cached input at $0.26. GLM-5.3-Flash costs $0.075 in and $0.25 out, and several older Flash models are priced at zero, a real free tier instead of trial credits. GLM-5, released in February 2026, is a 744B mixture-of-experts model under an MIT license, and at launch it ranked first among open-weight models on the Artificial Analysis index with a record-low hallucination score.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Z.ai for

- Budget coding inside Claude Code on the GLM plan
- A free Flash tier for prototyping
- Cheap GLM-5.3 inference at 1M context

### Choose Thinking Machines for

- RL or SFT on GLM-5.3 without owning GPUs
- US-based lab for teams avoiding China-hosted servers
- Multimodal Inkling with audio input

## At a glance

| Attribute | Z.ai | Thinking Machines |
|---|---|---|
| Model access | Open weights (MIT) | Open weights |
| Flagship models | GLM-5.3, GLM-5.3-Flash | Inkling, Inkling-Small |
| Speed | ~80 tok/s on GLM-5.3 | - |
| Price | $1.40 in, $4.40 out (GLM-5.3); free Flash tier | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Open weights, no license limits | LoRA SFT and RL via Tinker |
| Deployment | API, GLM Coding Plan | Training API, beta serverless (Inkling only) |
| Long context | 1M (GLM-5.3) | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Z.ai and Thinking Machines?

Z.ai sells GLM-5.3 cheaply with a flat-rate coding plan. Thinking Machines lets teams post-train GLM-5.3 with Tinker and offers its own Inkling models.

### When should I choose Z.ai over Thinking Machines?

Budget coding inside Claude Code on the GLM plan; A free Flash tier for prototyping; Cheap GLM-5.3 inference at 1M context.

### When should I choose Thinking Machines over Z.ai?

RL or SFT on GLM-5.3 without owning GPUs; US-based lab for teams avoiding China-hosted servers; Multimodal Inkling with audio input.

### Is Z.ai or Thinking Machines cheaper?

Z.ai: $1.40 in, $4.40 out (GLM-5.3); free Flash tier. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Z.ai or Thinking Machines?

Z.ai: 1M (GLM-5.3). Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Z.ai](https://www.subconscious.dev/compare/subconscious-vs-z-ai.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Z.ai](https://www.subconscious.dev/providers/z-ai.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
