Modal vs StepFun
StepFun ships cheap multimodal Step models, many Apache 2.0, from a China-hosted API. Modal gives you per-second GPUs to self-host those weights. API convenience versus control over where it runs.
By The Subconscious Team · Updated
Modal vs StepFun: key differences
StepFun is a Shanghai lab. Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active, 256K context, selectable reasoning and Apache 2.0 weights, priced at $0.20 in and $1.15 out on StepFun's API, which is China-hosted. The open weights run on vLLM and SGLang. Modal sells serverless GPU containers, up to 8 GPUs each across T4 through B300, and is a natural place to run those weights with your own serving code. Those weights are why the comparison exists at all.
Choose StepFun's API when the goal is cheap vision and video understanding with no infrastructure. Choose Modal to self-host the same weights, which keeps traffic off China-hosted inference and allows fine-tuning. The small active-parameter count is what makes self-hosting cheap, so a modest GPU configuration may be enough. Modal's cost risks are cold starts from loading weights and warm-container bills. StepFun's are thin Western support and models that trail the frontier on hard multimodal reasoning.
What Modal and StepFun do
Modal
Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.
Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper
Full Modal profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose Modal or StepFun?
Modal vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Bring your own weights | Open (Apache 2.0) and API models |
| Flagship models | None hosted | Step 3.7 Flash, Step3 |
| Speed | ~1s container boot | ~128 tok/s on Step 3.7 Flash |
| Price | Per second; H100 $3.95/hr list | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Run any training code | Open weights to fine-tune |
| Deployment | Serverless GPU containers | First-party API, OpenRouter |
| Long context | Depends on the model you deploy | 256K |
Frequently asked questions
What is the difference between Modal and StepFun?
StepFun ships cheap multimodal Step models, many Apache 2.0, from a China-hosted API. Modal gives you per-second GPUs to self-host those weights. API convenience versus control over where it runs.
When should I choose Modal over StepFun?
Self-hosting Step weights outside China; Fine-tuning Apache 2.0 models under your control; Bursty multimodal jobs billed by the second.
When should I choose StepFun over Modal?
Cheap hosted vision and video understanding; 256K context on a low-cost multimodal API; Prototyping with no GPU setup.
Is Modal or StepFun cheaper?
Modal: Per second; H100 $3.95/hr list. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, Modal or StepFun?
Modal: Depends on the model you deploy. StepFun: 256K.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.