vs

Relace vs Wafer

Wafer hosts big open models tuned for speed, sold through a flat pass for coding tools. Relace hosts small tool models for applying edits and searching code. They stack rather than compete.

By The Subconscious Team · Updated

Relace vs Wafer: key differences

Wafer speeds up general open models. Its agents tune the serving stack for each workload, and it reports a tuned Qwen 3.5 397B running 2.8x faster than stock SGLang, with GLM 5.1 and DeepSeek V4 Pro each 2x faster than a vLLM baseline, all self-reported. Wafer Pass, from $10 a week, covers every hosted model in Claude Code, Cline and OpenHands. Relace does not host general models. relace-apply-3 merges edits at about 10,000 tokens per second, and Relace adds agentic search and compaction at 50,000 tokens per second.

A coding agent can run a Wafer-hosted open model as its main brain and hand edits and retrieval to Relace, cutting work the big model would otherwise do. Wafer suits teams with a strict latency SLO who want a dedicated, continually retuned deployment on NVIDIA or AMD. Relace suits teams that want to self-host tool models next to their code. Both are young, specialized vendors. Wafer's catalog is small, and Relace's apply model stops at 128K tokens.

What Relace and Wafer do

Relace

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

Example models: relace-apply-3, Relace agentic search

Full Relace profile

Wafer

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo

Full Wafer profile

Should you choose Relace or Wafer?

Relace

Choose Relace for

  • Offloading file edits from the main model
  • Codebase search in seconds
  • Self-hosted utility models

Wafer

Choose Wafer for

  • Fast main models for coding agents on a flat pass
  • Dedicated endpoints tuned to a latency SLO
  • Serving on NVIDIA or AMD hardware

Relace vs Wafer at a glance

AttributeRelaceWafer
Model accessSpecialist modelsOpen weights
Flagship modelsrelace-apply-3, agentic searchQwen 3.5 397B Turbo, GLM 5.1 Turbo
Speed~10,000 tok/s apply2–2.8x vs stock vLLM or SGLang
Price3x+ cheaper than full rewritesWafer Pass from $10 a week
CustomizationUnknownAgent-tuned dedicated deployments
DeploymentHosted API or self-hostedServerless pass, dedicated
Long context128K maxVaries by model

Frequently asked questions

What is the difference between Relace and Wafer?

Wafer hosts big open models tuned for speed, sold through a flat pass for coding tools. Relace hosts small tool models for applying edits and searching code. They stack rather than compete.

When should I choose Relace over Wafer?

Offloading file edits from the main model; Codebase search in seconds; Self-hosted utility models.

When should I choose Wafer over Relace?

Fast main models for coding agents on a flat pass; Dedicated endpoints tuned to a latency SLO; Serving on NVIDIA or AMD hardware.

Is Relace or Wafer cheaper?

Relace: 3x+ cheaper than full rewrites. Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.

Which has more context, Relace or Wafer?

Relace: 128K max. Wafer: Varies by model.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.