# StepFun vs Wafer

> A lab that designs cheap-to-serve multimodal models against a startup that tunes serving stacks for open models. Both chase lower inference cost from different ends.

Canonical: https://www.subconscious.dev/compare/stepfun-vs-wafer · By The Subconscious Team · Updated September 30, 2026

## How they compare

StepFun and Wafer attack the same cost problem from opposite sides. StepFun works on the model. Its research co-designs models and systems to cut decoding cost, as with Step3's Multi-Matrix Factorization Attention and Attention-FFN Disaggregation, and Step 3.7 Flash keeps only 11B of its 198B parameters active. Wafer works on serving. Its agents profile a workload and tune batching, decoding, quantization, kernels and hardware, then keep re-tuning. Wafer reports GLM 5.1 and DeepSeek V4 Pro running 2x faster than a vLLM baseline, a self-reported figure.

As products, they rarely overlap. StepFun sells its own multimodal models through an API, at $0.20 in and $1.15 out for Step 3.7 Flash, plus Apache 2.0 weights. Wafer sells a small catalog of big open models through Wafer Pass from $10 a week, and dedicated deployments built around a customer's model and SLO on NVIDIA or AMD. A team self-hosting Step weights could in principle ask Wafer to tune that deployment. Both carry caveats: Wafer is very young, and StepFun's first-party inference is China-hosted.

## What each one does

### StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

### Wafer

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

## Which is best, and when

### Choose StepFun for

- Vision and video understanding at low per-token cost
- A small-active-parameter model for cheap self-hosting
- Speech and audio models from the same lab

### Choose Wafer for

- Big open models inside coding harnesses at a flat price
- Tuning a dedicated deployment to a latency SLO
- Serving across both NVIDIA and AMD

## At a glance

| Attribute | StepFun | Wafer |
|---|---|---|
| Model access | Open (Apache 2.0) and API models | Open weights |
| Flagship models | Step 3.7 Flash, Step3 | Qwen 3.5 397B Turbo, GLM 5.1 Turbo |
| Speed | ~128 tok/s on Step 3.7 Flash | 2–2.8x vs stock vLLM or SGLang |
| Price | $0.20 in, $1.15 out (Step 3.7 Flash) | Wafer Pass from $10 a week |
| Customization | Open weights to fine-tune | Agent-tuned dedicated deployments |
| Deployment | First-party API, OpenRouter | Serverless pass, dedicated |
| Long context | 256K | Varies by model |

## FAQ

### What is the difference between StepFun and Wafer?

A lab that designs cheap-to-serve multimodal models against a startup that tunes serving stacks for open models. Both chase lower inference cost from different ends.

### When should I choose StepFun over Wafer?

Vision and video understanding at low per-token cost; A small-active-parameter model for cheap self-hosting; Speech and audio models from the same lab.

### When should I choose Wafer over StepFun?

Big open models inside coding harnesses at a flat price; Tuning a dedicated deployment to a latency SLO; Serving across both NVIDIA and AMD.

### Is StepFun or Wafer cheaper?

StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.

### Which has more context, StepFun or Wafer?

StepFun: 256K. Wafer: Varies by model.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs StepFun](https://www.subconscious.dev/compare/subconscious-vs-stepfun.md), [Subconscious vs Wafer](https://www.subconscious.dev/compare/subconscious-vs-wafer.md).

Full profiles: [StepFun](https://www.subconscious.dev/providers/stepfun.md), [Wafer](https://www.subconscious.dev/providers/wafer.md).
