Cohere vs Wafer
Wafer tunes inference stacks so open models run faster on the same weights. Cohere ships its own Command models plus private, on-prem deployment.
By The Subconscious Team · Updated
Cohere vs Wafer: key differences
Wafer's product is speed on other labs' weights. Its agents profile a workload, test configs across batching, decoding, quantization, kernels and hardware, then keep re-tuning dedicated deployments on NVIDIA or AMD. Wafer reports its tuned Qwen 3.5 397B at 2.8x the speed of stock SGLang, and GLM 5.1 and DeepSeek V4 Pro at 2x a vLLM baseline. Wafer Pass is a flat-rate subscription from $10 a week that works in Claude Code, Cline and OpenHands. Cohere sells its own models by token, with Command A at $2.50 in and $10 out and 256K context, and reports 375 tokens per second on Command A+ in 4-bit form.
For coding agents, Wafer's catalog of large open models and its flat pricing are a better fit, since Command A+ trails the latest GLM and DeepSeek models on agentic coding. Cohere's strengths are elsewhere: Embed 4 and Rerank 4 for retrieval, Aya for multilingual use, managed fine-tuning, and deployment through Bedrock, Azure, OCI or fully on-prem. Wafer is a very young company with a small hosted catalog, and its speed claims are self-reported against stock baselines, so buyers should compare against tuned hosts. Enterprises with compliance reviews will find Cohere's track record easier to clear.
What Cohere and Wafer do
Cohere
Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.
Example models: Command A+, Command A, Embed 4, Rerank 4
Full Cohere profileWafer
Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.
Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo
Full Wafer profileShould you choose Cohere or Wafer?
Cohere vs Wafer at a glance
| Attribute | ||
|---|---|---|
| Model access | Closed, plus open Command A+ | Open weights |
| Flagship models | Command A+, Command A, Embed 4, Rerank 4 | Qwen 3.5 397B Turbo, GLM 5.1 Turbo |
| Speed | 375 tok/s on Command A+ W4A4, per Cohere | 2–2.8x vs stock vLLM or SGLang |
| Price | $0.0375–$2.50 in, $0.15–$10 out per 1M | Wafer Pass from $10 a week |
| Customization | Enterprise fine-tuning, incl. private | Agent-tuned dedicated deployments |
| Deployment | API, Bedrock, Azure, OCI, VPC, on-prem | Serverless pass, dedicated |
| Long context | 256K on Command A; 128K on A+ | Varies by model |
Frequently asked questions
What is the difference between Cohere and Wafer?
Wafer tunes inference stacks so open models run faster on the same weights. Cohere ships its own Command models plus private, on-prem deployment.
When should I choose Cohere over Wafer?
Compliance-reviewed enterprise deployments; Retrieval with Embed and Rerank; Fine-tuning inside a private network.
When should I choose Wafer over Cohere?
Big open models in coding harnesses; Flat weekly pricing for agent tools; Latency SLOs without in-house kernel work.
Is Cohere or Wafer cheaper?
Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.
Which has more context, Cohere or Wafer?
Cohere: 256K on Command A; 128K on A+. Wafer: Varies by model.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.