vs

Subconscious vs Wafer

Wafer tunes serving configurations on stock engines. Subconscious redesigns the runtime itself around long traces, with gains that grow past 200K tokens.

By The Subconscious Team · Updated

Subconscious vs Wafer: key differences

Wafer and Subconscious start from the same premise: stock vLLM or SGLang leaves performance on the table. Wafer's agents profile a workload, try configurations across batching, decoding, quantization, kernels and hardware, deploy the winner, then keep re-tuning. It reports Qwen 3.5 397B running 2.8x faster than stock SGLang, and GLM 5.1 and DeepSeek V4 Pro each 2x faster than a vLLM baseline. Subconscious changes the algorithm rather than the configuration. Its runtime drops in for vLLM or SGLang, prunes the KV cache and preserves suffix state, and delivers 2x faster task completion and 50% to 80% lower cost than standard inference, with gains that grow past 200K tokens. Wafer's figures are self-reported against stock baselines.

The pricing models differ sharply. Wafer Pass is a flat-rate subscription from $10 a week that covers every hosted model and drops into Claude Code, Cline and OpenHands, a good deal for individual developers. Subconscious bills processed tokens, which rewards long, cache-heavy traces at team scale. Wafer runs on NVIDIA and AMD, which helps teams hedge GPU supply, and it builds dedicated deployments around a customer's latency SLO. Subconscious adds Marathon post-trained model variants and on-prem options. Subconscious's dedicated deployments can also run nearly any open model.

What Subconscious and Wafer do

Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

Example models: GLM 5.3, DeepSeek V4.1 Flash

Full Subconscious profile

Wafer

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo

Full Wafer profile

Should you choose Subconscious or Wafer?

Subconscious

Choose Subconscious for

  • Agent traces past 200K tokens where cache pruning pays off
  • Team-scale billing on processed tokens
  • Marathon post-trained variants co-designed with the runtime

Wafer

Choose Wafer for

  • Individual developers who want flat-rate access in Claude Code
  • Dedicated endpoints tuned to a strict latency SLO
  • Hedging GPU supply across NVIDIA and AMD

Subconscious vs Wafer at a glance

AttributeSubconsciousWafer
Model accessOpen weightsOpen weights
Flagship modelsGLM 5.3, DeepSeek V4.1 FlashQwen 3.5 397B Turbo, GLM 5.1 Turbo
Speed2x faster task completion2–2.8x vs stock vLLM or SGLang
Price50–80% lower cost; billed on processed tokensWafer Pass from $10 a week
CustomizationMarathon post-trained variantsAgent-tuned dedicated deployments
DeploymentManaged API, dedicated, on-premServerless pass, dedicated
Long context5M+ effective contextVaries by model

Frequently asked questions

What is the difference between Subconscious and Wafer?

Wafer tunes serving configurations on stock engines. Subconscious redesigns the runtime itself around long traces, with gains that grow past 200K tokens.

When should I choose Subconscious over Wafer?

Agent traces past 200K tokens where cache pruning pays off; Team-scale billing on processed tokens; Marathon post-trained variants co-designed with the runtime.

When should I choose Wafer over Subconscious?

Individual developers who want flat-rate access in Claude Code; Dedicated endpoints tuned to a strict latency SLO; Hedging GPU supply across NVIDIA and AMD.

Is Subconscious or Wafer cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.

Which has more context, Subconscious or Wafer?

Subconscious: 5M+ effective context. Wafer: Varies by model.

Related comparisons

Run your longest agent traces on Subconscious

Point the OpenAI or Anthropic SDK, or the coding agent you already use, at Subconscious. Keep Wafer for the work it does best and send the long runs to us.