We raised $5.1M for long-running agents.
vs

Thinking Machines vs Wafer

Wafer tunes serving stacks so open models run faster on the same weights. Thinking Machines changes the weights themselves through LoRA post-training on Tinker.

By The Subconscious Team · Updated

Thinking Machines vs Wafer: key differences

Each works on a different layer. Wafer's agents profile a workload, try configs across batching, decoding, quantization, kernels and hardware, then deploy and keep re-tuning on NVIDIA or AMD. Wafer reports its Qwen 3.5 397B running 2.8x faster than stock SGLang, and GLM 5.1 and DeepSeek V4 Pro each 2x faster than a vLLM baseline. Its Wafer Pass starts at $10 a week for every hosted model and plugs into Claude Code, Cline and OpenHands. Thinking Machines leaves serving mostly alone. Tinker gives teams low-level training calls to run LoRA SFT or RL, with the lab running distributed GPU work.

They could work in sequence: train on Tinker, then serve on a Wafer dedicated deployment tuned to the traffic, if the base model and adapter are supported. Tinker's own sampling endpoint is scoped to testing and low internal traffic, and its serverless API covers only Inkling at $1.00 in and $4.05 out, far above Wafer's flat-rate pass for heavy agent use. Thinking Machines has much deeper funding and its own 975B Inkling with 1M context and audio input. Wafer is a very young company with a small hosted catalog, and its speedups are self-reported against stock baselines.

What Thinking Machines and Wafer do

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Wafer

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo

Full Wafer profile

Should you choose Thinking Machines or Wafer?

Thinking Machines

Choose Thinking Machines for

  • Changing model behavior with SFT or RL
  • LoRA training on large MoE bases
  • Inkling's audio input and 1M context

Wafer

Choose Wafer for

  • Faster open models on tuned serving stacks
  • Flat-rate access for coding agents
  • Dedicated endpoints with strict latency SLOs

Thinking Machines vs Wafer at a glance

AttributeThinking MachinesWafer
Model accessOpen weightsOpen weights
Flagship modelsInkling, Inkling-SmallQwen 3.5 397B Turbo, GLM 5.1 Turbo
SpeedUnknown2–2.8x vs stock vLLM or SGLang
PricePer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 outWafer Pass from $10 a week
CustomizationLoRA SFT and RL via TinkerAgent-tuned dedicated deployments
DeploymentTraining API, beta serverless (Inkling only)Serverless pass, dedicated
Long contextInkling up to 1M; Tinker 32K–256KVaries by model

Frequently asked questions

What is the difference between Thinking Machines and Wafer?

Wafer tunes serving stacks so open models run faster on the same weights. Thinking Machines changes the weights themselves through LoRA post-training on Tinker.

When should I choose Thinking Machines over Wafer?

Changing model behavior with SFT or RL; LoRA training on large MoE bases; Inkling's audio input and 1M context.

When should I choose Wafer over Thinking Machines?

Faster open models on tuned serving stacks; Flat-rate access for coding agents; Dedicated endpoints with strict latency SLOs.

Is Thinking Machines or Wafer cheaper?

Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.

Which has more context, Thinking Machines or Wafer?

Thinking Machines: Inkling up to 1M; Tinker 32K–256K. Wafer: Varies by model.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.