# Cohere vs Wafer

> Wafer tunes inference stacks so open models run faster on the same weights. Cohere ships its own Command models plus private, on-prem deployment.

Canonical: https://www.subconscious.dev/compare/cohere-vs-wafer · By The Subconscious Team · Updated September 30, 2026

## How they compare

Wafer's product is speed on other labs' weights. Its agents profile a workload, test configs across batching, decoding, quantization, kernels and hardware, then keep re-tuning dedicated deployments on NVIDIA or AMD. Wafer reports its tuned Qwen 3.5 397B at 2.8x the speed of stock SGLang, and GLM 5.1 and DeepSeek V4 Pro at 2x a vLLM baseline. Wafer Pass is a flat-rate subscription from $10 a week that works in Claude Code, Cline and OpenHands. Cohere sells its own models by token, with Command A at $2.50 in and $10 out and 256K context, and reports 375 tokens per second on Command A+ in 4-bit form.

For coding agents, Wafer's catalog of large open models and its flat pricing are a better fit, since Command A+ trails the latest GLM and DeepSeek models on agentic coding. Cohere's strengths are elsewhere: Embed 4 and Rerank 4 for retrieval, Aya for multilingual use, managed fine-tuning, and deployment through Bedrock, Azure, OCI or fully on-prem. Wafer is a very young company with a small hosted catalog, and its speed claims are self-reported against stock baselines, so buyers should compare against tuned hosts. Enterprises with compliance reviews will find Cohere's track record easier to clear.

## What each one does

### Cohere

Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.

### Wafer

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

## Which is best, and when

### Choose Cohere for

- Compliance-reviewed enterprise deployments
- Retrieval with Embed and Rerank
- Fine-tuning inside a private network

### Choose Wafer for

- Big open models in coding harnesses
- Flat weekly pricing for agent tools
- Latency SLOs without in-house kernel work

## At a glance

| Attribute | Cohere | Wafer |
|---|---|---|
| Model access | Closed, plus open Command A+ | Open weights |
| Flagship models | Command A+, Command A, Embed 4, Rerank 4 | Qwen 3.5 397B Turbo, GLM 5.1 Turbo |
| Speed | 375 tok/s on Command A+ W4A4, per Cohere | 2–2.8x vs stock vLLM or SGLang |
| Price | $0.0375–$2.50 in, $0.15–$10 out per 1M | Wafer Pass from $10 a week |
| Customization | Enterprise fine-tuning, incl. private | Agent-tuned dedicated deployments |
| Deployment | API, Bedrock, Azure, OCI, VPC, on-prem | Serverless pass, dedicated |
| Long context | 256K on Command A; 128K on A+ | Varies by model |

## FAQ

### What is the difference between Cohere and Wafer?

Wafer tunes inference stacks so open models run faster on the same weights. Cohere ships its own Command models plus private, on-prem deployment.

### When should I choose Cohere over Wafer?

Compliance-reviewed enterprise deployments; Retrieval with Embed and Rerank; Fine-tuning inside a private network.

### When should I choose Wafer over Cohere?

Big open models in coding harnesses; Flat weekly pricing for agent tools; Latency SLOs without in-house kernel work.

### Is Cohere or Wafer cheaper?

Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.

### Which has more context, Cohere or Wafer?

Cohere: 256K on Command A; 128K on A+. Wafer: Varies by model.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Cohere](https://www.subconscious.dev/compare/subconscious-vs-cohere.md), [Subconscious vs Wafer](https://www.subconscious.dev/compare/subconscious-vs-wafer.md).

Full profiles: [Cohere](https://www.subconscious.dev/providers/cohere.md), [Wafer](https://www.subconscious.dev/providers/wafer.md).
