vs

StepFun vs Wafer

A lab that designs cheap-to-serve multimodal models against a startup that tunes serving stacks for open models. Both chase lower inference cost from different ends.

By The Subconscious Team · Updated

StepFun vs Wafer: key differences

StepFun and Wafer attack the same cost problem from opposite sides. StepFun works on the model. Its research co-designs models and systems to cut decoding cost, as with Step3's Multi-Matrix Factorization Attention and Attention-FFN Disaggregation, and Step 3.7 Flash keeps only 11B of its 198B parameters active. Wafer works on serving. Its agents profile a workload and tune batching, decoding, quantization, kernels and hardware, then keep re-tuning. Wafer reports GLM 5.1 and DeepSeek V4 Pro running 2x faster than a vLLM baseline, a self-reported figure.

As products, they rarely overlap. StepFun sells its own multimodal models through an API, at $0.20 in and $1.15 out for Step 3.7 Flash, plus Apache 2.0 weights. Wafer sells a small catalog of big open models through Wafer Pass from $10 a week, and dedicated deployments built around a customer's model and SLO on NVIDIA or AMD. A team self-hosting Step weights could in principle ask Wafer to tune that deployment. Both carry caveats: Wafer is very young, and StepFun's first-party inference is China-hosted.

What StepFun and Wafer do

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Wafer

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo

Full Wafer profile

Should you choose StepFun or Wafer?

StepFun

Choose StepFun for

  • Vision and video understanding at low per-token cost
  • A small-active-parameter model for cheap self-hosting
  • Speech and audio models from the same lab

Wafer

Choose Wafer for

  • Big open models inside coding harnesses at a flat price
  • Tuning a dedicated deployment to a latency SLO
  • Serving across both NVIDIA and AMD

StepFun vs Wafer at a glance

AttributeStepFunWafer
Model accessOpen (Apache 2.0) and API modelsOpen weights
Flagship modelsStep 3.7 Flash, Step3Qwen 3.5 397B Turbo, GLM 5.1 Turbo
Speed~128 tok/s on Step 3.7 Flash2–2.8x vs stock vLLM or SGLang
Price$0.20 in, $1.15 out (Step 3.7 Flash)Wafer Pass from $10 a week
CustomizationOpen weights to fine-tuneAgent-tuned dedicated deployments
DeploymentFirst-party API, OpenRouterServerless pass, dedicated
Long context256KVaries by model

Frequently asked questions

What is the difference between StepFun and Wafer?

A lab that designs cheap-to-serve multimodal models against a startup that tunes serving stacks for open models. Both chase lower inference cost from different ends.

When should I choose StepFun over Wafer?

Vision and video understanding at low per-token cost; A small-active-parameter model for cheap self-hosting; Speech and audio models from the same lab.

When should I choose Wafer over StepFun?

Big open models inside coding harnesses at a flat price; Tuning a dedicated deployment to a latency SLO; Serving across both NVIDIA and AMD.

Is StepFun or Wafer cheaper?

StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.

Which has more context, StepFun or Wafer?

StepFun: 256K. Wafer: Varies by model.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.