vs

DeepInfra vs StepFun

StepFun is a Shanghai lab with efficient multimodal models, many under Apache 2.0. DeepInfra is a price-floor host for 150+ open models from many labs.

By The Subconscious Team · Updated

DeepInfra vs StepFun: key differences

StepFun builds models. DeepInfra hosts other people's. StepFun's workhorse, Step 3.7 Flash, is a 198B mixture-of-experts vision-language model with only 11B active parameters, 256K context, selectable reasoning levels and tool use. It ships under Apache 2.0, and StepFun's own API charges $0.20 in and $1.15 out per million. DeepInfra's comparable budget option would be something like DeepSeek V4 Flash at $0.14 in and $0.28 out, inside a catalog of 150+ models spanning text, image and speech. Whether DeepInfra lists StepFun's models is worth checking, since the open weights run on standard engines like vLLM and SGLang.

The choice mostly follows the workload and the location. StepFun's strength is low-cost vision and video understanding, and its small active parameter count makes self-hosting cheap. Its weaknesses are trailing frontier models on hard multimodal reasoning, thin Western distribution and support, and first-party inference hosted in China, though OpenRouter also carries Step 3.7 Flash. DeepInfra offers breadth and floor pricing on a fully OpenAI-compatible API, with quantization as the thing to check. Cost-sensitive agents that read images or video lean StepFun. Broad text workloads lean DeepInfra.

What DeepInfra and StepFun do

DeepInfra

DeepInfra is the price floor for open-model inference. Developers treat it as the reference point for what a token should cost, with small models like Llama 3.1 8B at $0.02 per million and DeepSeek V4 Flash at $0.14 in and $0.28 out. The catalog covers 150+ open models across text, image and speech behind a fully OpenAI-compatible API. There are no minimums, setup fees or contracts on the shared API.

Example models: DeepSeek V4 Flash, Llama 3.1 8B

Full DeepInfra profile

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Should you choose DeepInfra or StepFun?

DeepInfra

Choose DeepInfra for

  • Cheap text-heavy bulk jobs across many model families
  • Teams that want one key for dozens of labs
  • Broad text workloads on a fully OpenAI-compatible API

StepFun

Choose StepFun for

  • Vision and video understanding inside cost-sensitive agents
  • Self-hosting a small-active-parameter Apache 2.0 model
  • Tasks that need 256K context on a cheap multimodal model

DeepInfra vs StepFun at a glance

AttributeDeepInfraStepFun
Model accessOpen weightsOpen (Apache 2.0) and API models
Flagship modelsDeepSeek V4 Flash, Llama 3.1 8BStep 3.7 Flash, Step3
Speed~33 tok/s on DeepSeek V4 Pro (FP4)~128 tok/s on Step 3.7 Flash
PriceFrom $0.02 per 1M$0.20 in, $1.15 out (Step 3.7 Flash)
CustomizationNo managed fine-tuningOpen weights to fine-tune
DeploymentShared API, no contractsFirst-party API, OpenRouter
Long context66K on FP4 DeepSeek V4 Pro256K

Frequently asked questions

What is the difference between DeepInfra and StepFun?

StepFun is a Shanghai lab with efficient multimodal models, many under Apache 2.0. DeepInfra is a price-floor host for 150+ open models from many labs.

When should I choose DeepInfra over StepFun?

Cheap text-heavy bulk jobs across many model families; Teams that want one key for dozens of labs; Broad text workloads on a fully OpenAI-compatible API.

When should I choose StepFun over DeepInfra?

Vision and video understanding inside cost-sensitive agents; Self-hosting a small-active-parameter Apache 2.0 model; Tasks that need 256K context on a cheap multimodal model.

Is DeepInfra or StepFun cheaper?

DeepInfra: From $0.02 per 1M. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

Which has more context, DeepInfra or StepFun?

DeepInfra: 66K on FP4 DeepSeek V4 Pro. StepFun: 256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.