# StepFun vs RunInfra

> StepFun offers Apache 2.0 multimodal models at low token prices; RunInfra offers a small hosted library, $10 coding plans and an agent that builds endpoints.

Canonical: https://www.subconscious.dev/compare/stepfun-vs-runinfra · By The Subconscious Team · Updated September 30, 2026

## How they compare

Both providers push cheap models, but they package them differently. StepFun is a model lab. Its Step 3.7 Flash reads images and video with 256K context, tool use and structured outputs, at $0.20 in and $1.15 out, and its weights ship under Apache 2.0. RunInfra is a host with a tiny library of mid-size models such as Nemotron 3.5 Lightning 30B and Qwen 3.8 27B, sold per token with cached discounts or on coding plans from $10 a month that plug into Claude Code, Codex and many other CLIs.

RunInfra's deployment agent is its other half. It picks a model, benchmarks it across GPUs from L4 to B200, searches quantized variants like AWQ, GPTQ and FP8, and ships an endpoint that scales to zero, and paid plans accept custom uploads up to 50 GB. That could in principle host an open Step model, worth checking against the upload limit. For vision and video work, StepFun's own models are the direct path. For coding on a budget, RunInfra's plans fit. StepFun's first-party hosting is in China, and RunInfra is a young company.

## What each one does

### StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

### RunInfra

RunInfra pitches open models built for agents, with two ways in. Its hosted Model APIs serve a small curated library, including Nemotron 3.5 Lightning 30B, Qwen 3.8 27B and Ornith 1.5 35B, behind one key that works with both the OpenAI and Anthropic SDKs. Cached context bills at a discount. Coding plans start at $10 a month with limits that reset every five hours and every week, and they plug into Claude Code, Codex, OpenCode, Cline, Aider and dozens of other agent CLIs.

## Which is best, and when

### Choose StepFun for

- Multimodal understanding across images and video
- Apache 2.0 weights for self-hosting
- Structured outputs and tool use at low token prices

### Choose RunInfra for

- A $10 monthly coding plan for agent CLIs
- Automated quantization and GPU selection
- Scale-to-zero endpoints with cold starts under two seconds

## At a glance

| Attribute | StepFun | RunInfra |
|---|---|---|
| Model access | Open (Apache 2.0) and API models | Open weights |
| Flagship models | Step 3.7 Flash, Step3 | Nemotron 3.5 Lightning 30B, Qwen 3.8 27B |
| Speed | ~128 tok/s on Step 3.7 Flash | Cold starts under 2s |
| Price | $0.20 in, $1.15 out (Step 3.7 Flash) | Coding plans from $10 a month |
| Customization | Open weights to fine-tune | Uploads up to 50 GB; auto-quantization |
| Deployment | First-party API, OpenRouter | Model APIs, agent-built endpoints |
| Long context | 256K | Varies by model |

## FAQ

### What is the difference between StepFun and RunInfra?

StepFun offers Apache 2.0 multimodal models at low token prices; RunInfra offers a small hosted library, $10 coding plans and an agent that builds endpoints.

### When should I choose StepFun over RunInfra?

Multimodal understanding across images and video; Apache 2.0 weights for self-hosting; Structured outputs and tool use at low token prices.

### When should I choose RunInfra over StepFun?

A $10 monthly coding plan for agent CLIs; Automated quantization and GPU selection; Scale-to-zero endpoints with cold starts under two seconds.

### Is StepFun or RunInfra cheaper?

StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). RunInfra: Coding plans from $10 a month. The cheaper choice depends on the model and workload.

### Which has more context, StepFun or RunInfra?

StepFun: 256K. RunInfra: Varies by model.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs StepFun](https://www.subconscious.dev/compare/subconscious-vs-stepfun.md), [Subconscious vs RunInfra](https://www.subconscious.dev/compare/subconscious-vs-runinfra.md).

Full profiles: [StepFun](https://www.subconscious.dev/providers/stepfun.md), [RunInfra](https://www.subconscious.dev/providers/runinfra.md).
