DeepSeek vs StepFun
Two Chinese labs with permissive open weights. DeepSeek has stronger flagships and 1M context; StepFun's Step 3.7 Flash is a small-active-parameter multimodal model under Apache 2.0.
By The Subconscious Team · Updated
DeepSeek vs StepFun: key differences
DeepSeek and StepFun are both Chinese labs that pair a first-party API with open weights, DeepSeek under MIT and StepFun under Apache 2.0 on key releases. The models differ in shape. StepFun's Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and image and video understanding, at $0.20 in and $1.15 out. DeepSeek V4.1 Flash, at $0.30 in and $1.20 out at peak, adds image understanding with 1M context, and V4 Pro offers more capability at $1.32 in and $3.96 out. DeepSeek halves both off-peak.
StepFun trails frontier models on hard multimodal reasoning benchmarks, while DeepSeek's weights are among the strongest open models, which is why so many hosts serve them. StepFun's edge is video understanding and a small active parameter count that keeps self-hosting cheap. Both have China-hosted first-party inference, and StepFun's Western distribution and support are thin, though OpenRouter carries it. Pick StepFun for cheap vision and video inside agents. Pick DeepSeek for stronger reasoning, longer context and wider third-party hosting.
What DeepSeek and StepFun do
DeepSeek
DeepSeek is the Chinese lab whose open-weight models reset price expectations for the whole market. Its API now serves two models, both with 1M context and 384K max output. V4.1 Flash shipped September 10, 2026 with built-in image understanding at $0.30 in and $1.20 out at peak. V4 Pro, generally available since August 13, costs $1.32 in and $3.96 out at peak. Cache hits cost a few cents per million or less, and the weights ship on Hugging Face under an MIT license.
Example models: DeepSeek V4.1 Flash, DeepSeek V4 Pro
Full DeepSeek profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose DeepSeek or StepFun?
DeepSeek vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights (MIT) | Open (Apache 2.0) and API models |
| Flagship models | DeepSeek V4.1 Flash, V4 Pro | Step 3.7 Flash, Step3 |
| Speed | ~35 tok/s on V4 Pro | ~128 tok/s on Step 3.7 Flash |
| Price | Off-peak hours at half price | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Open weights to fine-tune | Open weights to fine-tune |
| Deployment | First-party API, Hugging Face weights | First-party API, OpenRouter |
| Long context | 1M, 384K max output | 256K |
Frequently asked questions
What is the difference between DeepSeek and StepFun?
Two Chinese labs with permissive open weights. DeepSeek has stronger flagships and 1M context; StepFun's Step 3.7 Flash is a small-active-parameter multimodal model under Apache 2.0.
When should I choose DeepSeek over StepFun?
Stronger reasoning on open weights; Contexts past 256K tokens; Wide choice of third-party hosts.
When should I choose StepFun over DeepSeek?
Video understanding at low cost; Cheap self-hosting with 11B active parameters; Apache 2.0 licensing.
Is DeepSeek or StepFun cheaper?
DeepSeek: Off-peak hours at half price. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, DeepSeek or StepFun?
DeepSeek: 1M, 384K max output. StepFun: 256K.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.