# Inference.net vs StepFun

> StepFun is a Shanghai lab with cheap Apache 2.0 multimodal models. Inference.net is a batch and distillation platform for open models.

Canonical: https://www.subconscious.dev/compare/inference-net-vs-stepfun · By The Subconscious Team · Updated September 30, 2026

## How they compare

StepFun builds models. Inference.net runs them. StepFun's Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and pricing of $0.20 in and $1.15 out on StepFun's API, released under Apache 2.0. Inference.net runs open models on aggregated spare GPU capacity through a Batch API with up to 1M requests per file, and its gateway routes to open, closed or custom models while capturing traffic for fine-tuning. StepFun's pitch is efficient multimodal models. Inference.net's pitch is cheap bulk compute and a loop from traces to a custom model.

StepFun's first-party inference is China-hosted, with thin Western distribution and support, which some buyers cannot accept. Its open weights, though, run on vLLM and SGLang wherever a team chooses. Inference.net focuses on cost through spare capacity and on custom distilled models, but offers few independent benchmarks. Cheap image and video understanding points to StepFun. Bulk text jobs, and turning production traces into a custom model, point to Inference.net.

## What each one does

### Inference.net

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

### StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

## Which is best, and when

### Choose Inference.net for

- Bulk text processing on discounted capacity
- Distilling a custom model from captured traffic
- Teams avoiding China-hosted first-party inference

### Choose StepFun for

- Cheap image and video understanding
- Self-hosting Apache 2.0 weights with small active parameters
- 256K context multimodal work

## At a glance

| Attribute | Inference.net | StepFun |
|---|---|---|
| Model access | Open, closed and custom | Open (Apache 2.0) and API models |
| Flagship models | Customer fine-tunes | Step 3.7 Flash, Step3 |
| Speed | Batch windows of 24h to 7 days | ~128 tok/s on Step 3.7 Flash |
| Price | Discounted spare GPU capacity | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Distill traces into custom models | Open weights to fine-tune |
| Deployment | Batch API, gateway, dedicated GPUs | First-party API, OpenRouter |
| Long context | Varies by model | 256K |

## FAQ

### What is the difference between Inference.net and StepFun?

StepFun is a Shanghai lab with cheap Apache 2.0 multimodal models. Inference.net is a batch and distillation platform for open models.

### When should I choose Inference.net over StepFun?

Bulk text processing on discounted capacity; Distilling a custom model from captured traffic; Teams avoiding China-hosted first-party inference.

### When should I choose StepFun over Inference.net?

Cheap image and video understanding; Self-hosting Apache 2.0 weights with small active parameters; 256K context multimodal work.

### Is Inference.net or StepFun cheaper?

Inference.net: Discounted spare GPU capacity. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

### Which has more context, Inference.net or StepFun?

Inference.net: Varies by model. StepFun: 256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Inference.net](https://www.subconscious.dev/compare/subconscious-vs-inference-net.md), [Subconscious vs StepFun](https://www.subconscious.dev/compare/subconscious-vs-stepfun.md).

Full profiles: [Inference.net](https://www.subconscious.dev/providers/inference-net.md), [StepFun](https://www.subconscious.dev/providers/stepfun.md).
