Thinking Machines vs Wafer
Wafer tunes serving stacks so open models run faster on the same weights. Thinking Machines changes the weights themselves through LoRA post-training on Tinker.
By The Subconscious Team · Updated
Thinking Machines vs Wafer: key differences
Each works on a different layer. Wafer's agents profile a workload, try configs across batching, decoding, quantization, kernels and hardware, then deploy and keep re-tuning on NVIDIA or AMD. Wafer reports its Qwen 3.5 397B running 2.8x faster than stock SGLang, and GLM 5.1 and DeepSeek V4 Pro each 2x faster than a vLLM baseline. Its Wafer Pass starts at $10 a week for every hosted model and plugs into Claude Code, Cline and OpenHands. Thinking Machines leaves serving mostly alone. Tinker gives teams low-level training calls to run LoRA SFT or RL, with the lab running distributed GPU work.
They could work in sequence: train on Tinker, then serve on a Wafer dedicated deployment tuned to the traffic, if the base model and adapter are supported. Tinker's own sampling endpoint is scoped to testing and low internal traffic, and its serverless API covers only Inkling at $1.00 in and $4.05 out, far above Wafer's flat-rate pass for heavy agent use. Thinking Machines has much deeper funding and its own 975B Inkling with 1M context and audio input. Wafer is a very young company with a small hosted catalog, and its speedups are self-reported against stock baselines.
What Thinking Machines and Wafer do
Thinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileWafer
Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.
Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo
Full Wafer profileShould you choose Thinking Machines or Wafer?
Thinking Machines
Choose Thinking Machines for
- Changing model behavior with SFT or RL
- LoRA training on large MoE bases
- Inkling's audio input and 1M context
Wafer
Choose Wafer for
- Faster open models on tuned serving stacks
- Flat-rate access for coding agents
- Dedicated endpoints with strict latency SLOs
Thinking Machines vs Wafer at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | Inkling, Inkling-Small | Qwen 3.5 397B Turbo, GLM 5.1 Turbo |
| Speed | Unknown | 2–2.8x vs stock vLLM or SGLang |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | Wafer Pass from $10 a week |
| Customization | LoRA SFT and RL via Tinker | Agent-tuned dedicated deployments |
| Deployment | Training API, beta serverless (Inkling only) | Serverless pass, dedicated |
| Long context | Inkling up to 1M; Tinker 32K–256K | Varies by model |
Frequently asked questions
What is the difference between Thinking Machines and Wafer?
Wafer tunes serving stacks so open models run faster on the same weights. Thinking Machines changes the weights themselves through LoRA post-training on Tinker.
When should I choose Thinking Machines over Wafer?
Changing model behavior with SFT or RL; LoRA training on large MoE bases; Inkling's audio input and 1M context.
When should I choose Wafer over Thinking Machines?
Faster open models on tuned serving stacks; Flat-rate access for coding agents; Dedicated endpoints with strict latency SLOs.
Is Thinking Machines or Wafer cheaper?
Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.
Which has more context, Thinking Machines or Wafer?
Thinking Machines: Inkling up to 1M; Tinker 32K–256K. Wafer: Varies by model.
Related comparisons
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Subconscious vs Wafer
OpenAI vs Wafer
Anthropic vs Wafer
Google Vertex AI vs Wafer
Amazon Bedrock vs Wafer
Together AI vs Wafer
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.