Modal vs Wafer
Wafer sells agent-tuned inference on open models via a flat pass or dedicated endpoints. Modal sells raw serverless GPUs. Tuned serving handed to you versus serving you build.
By The Subconscious Team · Updated
Modal vs Wafer: key differences
Wafer is what a team gets if it outsources the serving engineering. Its agents profile a workload, try configs across batching, decoding, quantization, engines, kernels and hardware, then deploy the winner and keep re-tuning, on NVIDIA or AMD. Wafer reports its Qwen 3.5 397B running 2.8x faster than stock SGLang. Modal gives developers the GPUs and container orchestration but leaves the serving stack to them, typically something like vLLM, which is the baseline Wafer says it beats. Wafer owns the tuning; on Modal, you do.
For coding agents on large open models, Wafer Pass from $10 a week is far simpler than running a 397B model on Modal. For a dedicated endpoint with a strict latency SLO and no kernel engineers, Wafer's managed tuning is the pitch. Modal wins for custom models, non-LLM work, fine-tuning and agent sandboxes, with a Python-native workflow and $30 in free monthly credits. Keep in mind that Wafer is a 2025 company with few hosted models, and its speed figures compare against untuned stacks.
What Modal and Wafer do
Modal
Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.
Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper
Full Modal profileWafer
Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.
Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo
Full Wafer profileShould you choose Modal or Wafer?
Modal
Choose Modal for
- Custom models and non-LLM GPU jobs.
- Teams that want to own their serving code.
- Fine-tuning, batch and sandboxes on one platform.
Wafer
Choose Wafer for
- Big open models at interactive speed for coding agents.
- Latency-SLO endpoints without in-house kernel engineers.
- Flat weekly pricing from $10.
Modal vs Wafer at a glance
| Attribute | ||
|---|---|---|
| Model access | Bring your own weights | Open weights |
| Flagship models | None hosted | Qwen 3.5 397B Turbo, GLM 5.1 Turbo |
| Speed | ~1s container boot | 2–2.8x vs stock vLLM or SGLang |
| Price | Per second; H100 $3.95/hr list | Wafer Pass from $10 a week |
| Customization | Run any training code | Agent-tuned dedicated deployments |
| Deployment | Serverless GPU containers | Serverless pass, dedicated |
| Long context | Depends on the model you deploy | Varies by model |
Frequently asked questions
What is the difference between Modal and Wafer?
Wafer sells agent-tuned inference on open models via a flat pass or dedicated endpoints. Modal sells raw serverless GPUs. Tuned serving handed to you versus serving you build.
When should I choose Modal over Wafer?
Custom models and non-LLM GPU jobs; Teams that want to own their serving code; Fine-tuning, batch and sandboxes on one platform.
When should I choose Wafer over Modal?
Big open models at interactive speed for coding agents; Latency-SLO endpoints without in-house kernel engineers; Flat weekly pricing from $10.
Is Modal or Wafer cheaper?
Modal: Per second; H100 $3.95/hr list. Wafer: Wafer Pass from $10 a week. The cheaper choice depends on the model and workload.
Which has more context, Modal or Wafer?
Modal: Depends on the model you deploy. Wafer: Varies by model.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.