vs

StepFun vs Runware

StepFun's models read images and video; Runware's generate them. A multimodal understanding lab and a media generation API cover opposite halves of the task.

By The Subconscious Team · Updated

StepFun vs Runware: key differences

StepFun and Runware both work with images and video, from opposite ends. Step 3.7 Flash is a vision-language model that reads images and video with 256K context, at $0.20 in and $1.15 out, and StepFun also builds speech and audio models. Runware generates media: image, video, audio and 3D through one request schema, with images from fractions of a cent and Seedance 2.5 video at about $0.10 a second at 480p. Runware carries text too, but LLM hosting is a side line next to its catalog of 300+ priced media models.

Together they form a loop. Runware renders a batch of images or clips, and Step 3.7 Flash reviews them, captions them or checks them against a brief, using tool use and structured outputs to return results in a fixed shape. Both sell on cost. StepFun keeps decoding cheap through model and system co-design and ships Apache 2.0 weights, while Runware says its Sonic Inference Engine prices around 10x lower. StepFun's first-party inference is China-hosted, and Runware's output URLs expire after seven days by default.

What StepFun and Runware do

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Runware

Runware sells what it calls the lowest-cost API for media generation, and it claims more than 1M developers. One endpoint covers image, video, audio, 3D and text. Every request is a task with the same shape, so switching from a Kling video to a Seedream image mostly means changing the model ID. The published rate sheet lists 300+ priced models, with images from fractions of a cent to a few cents each and video billed per second, like Seedance 2.5 at about $0.10 a second at 480p.

Example models: Seedance 2.5, Qwen-Image-3.0

Full Runware profile

Should you choose StepFun or Runware?

StepFun

Choose StepFun for

  • Captioning and checking generated images and video
  • Video understanding inside cost-sensitive agents
  • Self-hosting a multimodal model under Apache 2.0

Runware

Choose Runware for

  • Generating images, clips, audio and 3D at volume
  • Switching generation models by model ID
  • Renting H100s at $2.76 an hour for diffusion checkpoints

StepFun vs Runware at a glance

AttributeStepFunRunware
Model accessOpen (Apache 2.0) and API modelsHosted media models
Flagship modelsStep 3.7 Flash, Step3Seedance 2.5, Qwen-Image-3.0
Speed~128 tok/s on Step 3.7 FlashUnknown
Price$0.20 in, $1.15 out (Step 3.7 Flash)Images from fractions of a cent
CustomizationOpen weights to fine-tuneFine-tuned diffusion checkpoints
DeploymentFirst-party API, OpenRouterUnified API, raw GPUs
Long context256KNot applicable

Frequently asked questions

What is the difference between StepFun and Runware?

StepFun's models read images and video; Runware's generate them. A multimodal understanding lab and a media generation API cover opposite halves of the task.

When should I choose StepFun over Runware?

Captioning and checking generated images and video; Video understanding inside cost-sensitive agents; Self-hosting a multimodal model under Apache 2.0.

When should I choose Runware over StepFun?

Generating images, clips, audio and 3D at volume; Switching generation models by model ID; Renting H100s at $2.76 an hour for diffusion checkpoints.

Is StepFun or Runware cheaper?

StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). Runware: Images from fractions of a cent. The cheaper choice depends on the model and workload.

Which has more context, StepFun or Runware?

StepFun: 256K. Runware: Not applicable.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.