Relace vs Wafer
Wafer hosts big open models tuned for speed, sold through a flat pass for coding tools. Relace hosts small tool models for applying edits and searching code. They stack rather than compete.
By The Subconscious Team · Updated
Relace vs Wafer: key differences
Wafer speeds up general open models. Its agents tune the serving stack for each workload, and it reports a tuned Qwen 3.5 397B running 2.8x faster than stock SGLang, with GLM 5.1 and DeepSeek V4 Pro each 2x faster than a vLLM baseline, all self-reported. Wafer Pass, from $10 a week, covers every hosted model in Claude Code, Cline and OpenHands. Relace does not host general models. relace-apply-3 merges edits at about 10,000 tokens per second, and Relace adds agentic search and compaction at 50,000 tokens per second.
A coding agent can run a Wafer-hosted open model as its main brain and hand edits and retrieval to Relace, cutting work the big model would otherwise do. Wafer suits teams with a strict latency SLO who want a dedicated, continually retuned deployment on NVIDIA or AMD. Relace suits teams that want to self-host tool models next to their code. Both are young, specialized vendors. Wafer's catalog is small, and Relace's apply model stops at 128K tokens.
What Relace and Wafer do
Relace
Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.
Example models: relace-apply-3, Relace agentic search
Full Relace profileWafer
Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.
Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo
Full Wafer profileShould you choose Relace or Wafer?
Relace vs Wafer at a glance
| Attribute | ||
|---|---|---|
| Model access | Specialist models | Open weights |
| Flagship models | relace-apply-3, agentic search | Qwen 3.5 397B Turbo, GLM 5.1 Turbo |
| Speed | ~10,000 tok/s apply | 2–2.8x vs stock vLLM or SGLang |
| Price | 3x+ cheaper than full rewrites | Wafer Pass from $10 a week |
| Customization | Unknown | Agent-tuned dedicated deployments |
| Deployment | Hosted API or self-hosted | Serverless pass, dedicated |
| Long context | 128K max | Varies by model |
Frequently asked questions
What is the difference between Relace and Wafer?
Wafer hosts big open models tuned for speed, sold through a flat pass for coding tools. Relace hosts small tool models for applying edits and searching code. They stack rather than compete.
When should I choose Relace over Wafer?
Offloading file edits from the main model; Codebase search in seconds; Self-hosted utility models.
When should I choose Wafer over Relace?
Fast main models for coding agents on a flat pass; Dedicated endpoints tuned to a latency SLO; Serving on NVIDIA or AMD hardware.
Is Relace or Wafer cheaper?
Relace: 3x+ cheaper than full rewrites. Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.
Which has more context, Relace or Wafer?
Relace: 128K max. Wafer: Varies by model.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.