# Wafer

> Agent-tuned inference stacks that run open models faster on the same weights.

Canonical: https://www.subconscious.dev/providers/wafer · By The Subconscious Team · Updated September 30, 2026

- Founded: 2025
- Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo
- Website: https://www.wafer.ai

## Overview

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

Wafer calls the result continual inference. A dedicated deployment gets built around a customer's model, traffic shape and SLO, and the agent keeps re-tuning as load, models or hardware change, on NVIDIA or AMD. The serverless side, Wafer Pass, is a flat-rate subscription from $10 a week that covers every hosted model and drops into Claude Code, Cline, OpenHands and similar harnesses. Wafer reports its tuned Qwen 3.5 397B running 2.8x faster than stock SGLang, with GLM 5.1 and DeepSeek V4 Pro each 2x faster than a vLLM baseline.

## Upsides

- Faster open models on the same weights, since most hosts run stock vLLM or SGLang.
- Cheap flat-rate access for agentic coding tools.
- Works across NVIDIA and AMD, which helps teams hedge GPU supply.

## Core use cases

- Coding agents that want big open models at interactive speed.
- Dedicated endpoints for teams with a strict latency SLO and no in-house kernel engineers.

## Downsides

- Very young company with a small hosted catalog.
- Speedup claims are self-reported against stock baselines, so compare against tuned hosts like Fireworks before buying.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open weights |
| Flagship models | Qwen 3.5 397B Turbo, GLM 5.1 Turbo |
| Speed | 2–2.8x vs stock vLLM or SGLang |
| Price | Wafer Pass from $10 a week |
| Customization | Agent-tuned dedicated deployments |
| Deployment | Serverless pass, dedicated |
| Long context | Varies by model |

## FAQ

### What is Wafer?

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

### What is Wafer best for?

Coding agents that want big open models at interactive speed; Dedicated endpoints for teams with a strict latency SLO and no in-house kernel engineers.

### How much does Wafer cost?

Wafer pricing at a glance: Wafer Pass from $10 a week. Rates change often, so check Wafer's pricing page before committing.

### How much context does Wafer support?

Wafer's long-context support: Varies by model.

### What are the downsides of Wafer?

Very young company with a small hosted catalog; Speedup claims are self-reported against stock baselines, so compare against tuned hosts like Fireworks before buying.

### What are the best alternatives to Wafer?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Wafer on this site.

## Comparisons

- [Subconscious vs Wafer](https://www.subconscious.dev/compare/subconscious-vs-wafer.md)
- [OpenAI vs Wafer](https://www.subconscious.dev/compare/openai-vs-wafer.md)
- [Anthropic vs Wafer](https://www.subconscious.dev/compare/anthropic-vs-wafer.md)
- [Google Vertex AI vs Wafer](https://www.subconscious.dev/compare/google-vertex-vs-wafer.md)
- [Amazon Bedrock vs Wafer](https://www.subconscious.dev/compare/aws-bedrock-vs-wafer.md)
- [Together AI vs Wafer](https://www.subconscious.dev/compare/together-ai-vs-wafer.md)
- [Fireworks AI vs Wafer](https://www.subconscious.dev/compare/fireworks-vs-wafer.md)
- [Baseten vs Wafer](https://www.subconscious.dev/compare/baseten-vs-wafer.md)
- [Groq vs Wafer](https://www.subconscious.dev/compare/groq-vs-wafer.md)
- [Cerebras vs Wafer](https://www.subconscious.dev/compare/cerebras-vs-wafer.md)
- [DeepInfra vs Wafer](https://www.subconscious.dev/compare/deepinfra-vs-wafer.md)
- [Modal vs Wafer](https://www.subconscious.dev/compare/modal-vs-wafer.md)
- [xAI vs Wafer](https://www.subconscious.dev/compare/xai-vs-wafer.md)
- [DeepSeek vs Wafer](https://www.subconscious.dev/compare/deepseek-vs-wafer.md)
- [Moonshot AI vs Wafer](https://www.subconscious.dev/compare/moonshot-ai-vs-wafer.md)
- [Z.ai vs Wafer](https://www.subconscious.dev/compare/z-ai-vs-wafer.md)
- [Alibaba Cloud vs Wafer](https://www.subconscious.dev/compare/alibaba-cloud-vs-wafer.md)
- [Meta vs Wafer](https://www.subconscious.dev/compare/meta-vs-wafer.md)
- [SambaNova vs Wafer](https://www.subconscious.dev/compare/sambanova-vs-wafer.md)
- [Nebius vs Wafer](https://www.subconscious.dev/compare/nebius-vs-wafer.md)
- [fal vs Wafer](https://www.subconscious.dev/compare/fal-vs-wafer.md)
- [Novita AI vs Wafer](https://www.subconscious.dev/compare/novita-ai-vs-wafer.md)
- [Parasail vs Wafer](https://www.subconscious.dev/compare/parasail-vs-wafer.md)
- [Inference.net vs Wafer](https://www.subconscious.dev/compare/inference-net-vs-wafer.md)
- [GMI Cloud vs Wafer](https://www.subconscious.dev/compare/gmi-cloud-vs-wafer.md)
- [Sail Research vs Wafer](https://www.subconscious.dev/compare/sail-research-vs-wafer.md)
- [Morph vs Wafer](https://www.subconscious.dev/compare/morph-vs-wafer.md)
- [Relace vs Wafer](https://www.subconscious.dev/compare/relace-vs-wafer.md)
- [TypeSafe AI vs Wafer](https://www.subconscious.dev/compare/typesafe-ai-vs-wafer.md)
- [StepFun vs Wafer](https://www.subconscious.dev/compare/stepfun-vs-wafer.md)
- [Runware vs Wafer](https://www.subconscious.dev/compare/runware-vs-wafer.md)
- [StreamLake vs Wafer](https://www.subconscious.dev/compare/streamlake-vs-wafer.md)
- [Wafer vs RunInfra](https://www.subconscious.dev/compare/wafer-vs-runinfra.md)
- [Wafer vs Particle.AI](https://www.subconscious.dev/compare/wafer-vs-particle-ai.md)

## Sources

- [Wafer](https://www.wafer.ai/)
- [Wafer company info](https://www.wafer.ai/ai-info)
- [Wafer on Y Combinator](https://www.ycombinator.com/companies/wafer)

Pricing and model lineups change often; figures are a snapshot.
