Baseten vs StepFun
StepFun is a Shanghai lab with cheap, efficient multimodal models, many under Apache 2.0. Baseten is a host that could serve those weights outside China.
By The Subconscious Team · Updated
Baseten vs StepFun: key differences
StepFun builds models; Baseten serves them. StepFun's workhorse, Step 3.7 Flash, is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and an Apache 2.0 license, priced at $0.20 in and $1.15 out on StepFun's own API. That is cheap for a model that reads images and video. The weak spots are distribution and location: StepFun's first-party inference is China-hosted, and Western support is thin. Step 3.7 Flash is not in Baseten's 13-model list, but its open weights could run on a Baseten dedicated deployment through Truss.
That route matters mainly for companies that want StepFun's efficiency without sending data to China. Baseten brings HIPAA, data residency, a 99.99% SLA and the lowest measured time to first token, at the cost of per-minute GPU billing (an H100 at about $6.50 an hour) instead of StepFun's per-token rates. StepFun's own benchmarks trail frontier models on hard multimodal reasoning. For a cost-sensitive vision agent with no residency concerns, calling StepFun directly is simpler.
What Baseten and StepFun do
Baseten
Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.
Example models: GLM 5.2, gpt-oss 120B
Full Baseten profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose Baseten or StepFun?
Baseten
Choose Baseten for
- Serving Apache 2.0 Step weights with data residency
- Curated text models like DeepSeek V4 and GLM 5.2
- Latency-bound production under a 99.99% SLA
StepFun
Choose StepFun for
- Cheap image and video understanding at 256K context
- Pay-per-token access with no GPU setup
- Small-active-parameter models for self-hosting
Baseten vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights, 13 curated | Open (Apache 2.0) and API models |
| Flagship models | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | Step 3.7 Flash, Step3 |
| Speed | 0.49s TTFT, lowest measured | ~128 tok/s on Step 3.7 Flash |
| Price | H100 about $6.50/hr dedicated | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Deploy any model with Truss | Open weights to fine-tune |
| Deployment | Model APIs, dedicated, self-host | First-party API, OpenRouter |
| Long context | Varies by model | 256K |
Frequently asked questions
What is the difference between Baseten and StepFun?
StepFun is a Shanghai lab with cheap, efficient multimodal models, many under Apache 2.0. Baseten is a host that could serve those weights outside China.
When should I choose Baseten over StepFun?
Serving Apache 2.0 Step weights with data residency; Curated text models like DeepSeek V4 and GLM 5.2; Latency-bound production under a 99.99% SLA.
When should I choose StepFun over Baseten?
Cheap image and video understanding at 256K context; Pay-per-token access with no GPU setup; Small-active-parameter models for self-hosting.
Is Baseten or StepFun cheaper?
Baseten: H100 about $6.50/hr dedicated. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, Baseten or StepFun?
Baseten: Varies by model. StepFun: 256K.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.