vs

Inference.net vs StepFun

StepFun is a Shanghai lab with cheap Apache 2.0 multimodal models. Inference.net is a batch and distillation platform for open models.

By The Subconscious Team · Updated

Inference.net vs StepFun: key differences

StepFun builds models. Inference.net runs them. StepFun's Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and pricing of $0.20 in and $1.15 out on StepFun's API, released under Apache 2.0. Inference.net runs open models on aggregated spare GPU capacity through a Batch API with up to 1M requests per file, and its gateway routes to open, closed or custom models while capturing traffic for fine-tuning. StepFun's pitch is efficient multimodal models. Inference.net's pitch is cheap bulk compute and a loop from traces to a custom model.

StepFun's first-party inference is China-hosted, with thin Western distribution and support, which some buyers cannot accept. Its open weights, though, run on vLLM and SGLang wherever a team chooses. Inference.net focuses on cost through spare capacity and on custom distilled models, but offers few independent benchmarks. Cheap image and video understanding points to StepFun. Bulk text jobs, and turning production traces into a custom model, point to Inference.net.

What Inference.net and StepFun do

Inference.net

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

Example models: open catalog models plus customer fine-tunes served on dedicated GPUs

Full Inference.net profile

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Should you choose Inference.net or StepFun?

Inference.net

Choose Inference.net for

  • Bulk text processing on discounted capacity
  • Distilling a custom model from captured traffic
  • Teams avoiding China-hosted first-party inference

StepFun

Choose StepFun for

  • Cheap image and video understanding
  • Self-hosting Apache 2.0 weights with small active parameters
  • 256K context multimodal work

Inference.net vs StepFun at a glance

AttributeInference.netStepFun
Model accessOpen, closed and customOpen (Apache 2.0) and API models
Flagship modelsCustomer fine-tunesStep 3.7 Flash, Step3
SpeedBatch windows of 24h to 7 days~128 tok/s on Step 3.7 Flash
PriceDiscounted spare GPU capacity$0.20 in, $1.15 out (Step 3.7 Flash)
CustomizationDistill traces into custom modelsOpen weights to fine-tune
DeploymentBatch API, gateway, dedicated GPUsFirst-party API, OpenRouter
Long contextVaries by model256K

Frequently asked questions

What is the difference between Inference.net and StepFun?

StepFun is a Shanghai lab with cheap Apache 2.0 multimodal models. Inference.net is a batch and distillation platform for open models.

When should I choose Inference.net over StepFun?

Bulk text processing on discounted capacity; Distilling a custom model from captured traffic; Teams avoiding China-hosted first-party inference.

When should I choose StepFun over Inference.net?

Cheap image and video understanding; Self-hosting Apache 2.0 weights with small active parameters; 256K context multimodal work.

Is Inference.net or StepFun cheaper?

Inference.net: Discounted spare GPU capacity. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

Which has more context, Inference.net or StepFun?

Inference.net: Varies by model. StepFun: 256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.