# Subconscious vs Wafer

> Wafer tunes serving configurations on stock engines. Subconscious redesigns the runtime itself around long traces, with gains that grow past 200K tokens.

Canonical: https://www.subconscious.dev/compare/subconscious-vs-wafer · By The Subconscious Team · Updated September 30, 2026

## How they compare

Wafer and Subconscious start from the same premise: stock vLLM or SGLang leaves performance on the table. Wafer's agents profile a workload, try configurations across batching, decoding, quantization, kernels and hardware, deploy the winner, then keep re-tuning. It reports Qwen 3.5 397B running 2.8x faster than stock SGLang, and GLM 5.1 and DeepSeek V4 Pro each 2x faster than a vLLM baseline. Subconscious changes the algorithm rather than the configuration. Its runtime drops in for vLLM or SGLang, prunes the KV cache and preserves suffix state, and delivers 2x faster task completion and 50% to 80% lower cost than standard inference, with gains that grow past 200K tokens. Wafer's figures are self-reported against stock baselines.

The pricing models differ sharply. Wafer Pass is a flat-rate subscription from $10 a week that covers every hosted model and drops into Claude Code, Cline and OpenHands, a good deal for individual developers. Subconscious bills processed tokens, which rewards long, cache-heavy traces at team scale. Wafer runs on NVIDIA and AMD, which helps teams hedge GPU supply, and it builds dedicated deployments around a customer's latency SLO. Subconscious adds Marathon post-trained model variants and on-prem options. Subconscious's dedicated deployments can also run nearly any open model.

## What each one does

### Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### Wafer

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

## Which is best, and when

### Choose Subconscious for

- Agent traces past 200K tokens where cache pruning pays off
- Team-scale billing on processed tokens
- Marathon post-trained variants co-designed with the runtime

### Choose Wafer for

- Individual developers who want flat-rate access in Claude Code
- Dedicated endpoints tuned to a strict latency SLO
- Hedging GPU supply across NVIDIA and AMD

## At a glance

| Attribute | Subconscious | Wafer |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | Qwen 3.5 397B Turbo, GLM 5.1 Turbo |
| Speed | 2x faster task completion | 2–2.8x vs stock vLLM or SGLang |
| Price | 50–80% lower cost; billed on processed tokens | Wafer Pass from $10 a week |
| Customization | Marathon post-trained variants | Agent-tuned dedicated deployments |
| Deployment | Managed API, dedicated, on-prem | Serverless pass, dedicated |
| Long context | 5M+ effective context | Varies by model |

## FAQ

### What is the difference between Subconscious and Wafer?

Wafer tunes serving configurations on stock engines. Subconscious redesigns the runtime itself around long traces, with gains that grow past 200K tokens.

### When should I choose Subconscious over Wafer?

Agent traces past 200K tokens where cache pruning pays off; Team-scale billing on processed tokens; Marathon post-trained variants co-designed with the runtime.

### When should I choose Wafer over Subconscious?

Individual developers who want flat-rate access in Claude Code; Dedicated endpoints tuned to a strict latency SLO; Hedging GPU supply across NVIDIA and AMD.

### Is Subconscious or Wafer cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.

### Which has more context, Subconscious or Wafer?

Subconscious: 5M+ effective context. Wafer: Varies by model.

## Try Subconscious

Subconscious speaks the OpenAI and Anthropic API formats. Base URL: https://api.subconscious.dev/v1. Docs: https://docs.subconscious.dev. Get an API key: https://platform.subconscious.dev/signin. Agent guide: https://www.subconscious.dev/agents.md.

Full profiles: [Subconscious](https://www.subconscious.dev/providers/subconscious.md), [Wafer](https://www.subconscious.dev/providers/wafer.md).
