Inference.net vs StepFun
StepFun is a Shanghai lab with cheap Apache 2.0 multimodal models. Inference.net is a batch and distillation platform for open models.
By The Subconscious Team · Updated
Inference.net vs StepFun: key differences
StepFun builds models. Inference.net runs them. StepFun's Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and pricing of $0.20 in and $1.15 out on StepFun's API, released under Apache 2.0. Inference.net runs open models on aggregated spare GPU capacity through a Batch API with up to 1M requests per file, and its gateway routes to open, closed or custom models while capturing traffic for fine-tuning. StepFun's pitch is efficient multimodal models. Inference.net's pitch is cheap bulk compute and a loop from traces to a custom model.
StepFun's first-party inference is China-hosted, with thin Western distribution and support, which some buyers cannot accept. Its open weights, though, run on vLLM and SGLang wherever a team chooses. Inference.net focuses on cost through spare capacity and on custom distilled models, but offers few independent benchmarks. Cheap image and video understanding points to StepFun. Bulk text jobs, and turning production traces into a custom model, point to Inference.net.
What Inference.net and StepFun do
Inference.net
Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.
Example models: open catalog models plus customer fine-tunes served on dedicated GPUs
Full Inference.net profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose Inference.net or StepFun?
Inference.net
Choose Inference.net for
- Bulk text processing on discounted capacity
- Distilling a custom model from captured traffic
- Teams avoiding China-hosted first-party inference
StepFun
Choose StepFun for
- Cheap image and video understanding
- Self-hosting Apache 2.0 weights with small active parameters
- 256K context multimodal work
Inference.net vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Open, closed and custom | Open (Apache 2.0) and API models |
| Flagship models | Customer fine-tunes | Step 3.7 Flash, Step3 |
| Speed | Batch windows of 24h to 7 days | ~128 tok/s on Step 3.7 Flash |
| Price | Discounted spare GPU capacity | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Distill traces into custom models | Open weights to fine-tune |
| Deployment | Batch API, gateway, dedicated GPUs | First-party API, OpenRouter |
| Long context | Varies by model | 256K |
Frequently asked questions
What is the difference between Inference.net and StepFun?
StepFun is a Shanghai lab with cheap Apache 2.0 multimodal models. Inference.net is a batch and distillation platform for open models.
When should I choose Inference.net over StepFun?
Bulk text processing on discounted capacity; Distilling a custom model from captured traffic; Teams avoiding China-hosted first-party inference.
When should I choose StepFun over Inference.net?
Cheap image and video understanding; Self-hosting Apache 2.0 weights with small active parameters; 256K context multimodal work.
Is Inference.net or StepFun cheaper?
Inference.net: Discounted spare GPU capacity. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, Inference.net or StepFun?
Inference.net: Varies by model. StepFun: 256K.
Related comparisons
Subconscious vs Inference.net
OpenAI vs Inference.net
Anthropic vs Inference.net
Google Vertex AI vs Inference.net
Amazon Bedrock vs Inference.net
Together AI vs Inference.net
Subconscious vs StepFun
OpenAI vs StepFun
Anthropic vs StepFun
Google Vertex AI vs StepFun
Amazon Bedrock vs StepFun
Together AI vs StepFun
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.