Cerebras vs StepFun
StepFun ships cheap Apache 2.0 multimodal models from Shanghai. Cerebras serves a small set of text models at the fastest public speeds.
By The Subconscious Team · Updated
Cerebras vs StepFun: key differences
StepFun is a lab and Cerebras is a host, and they barely overlap. StepFun's Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and selectable reasoning, priced at $0.20 in and $1.15 out on its own API and released under Apache 2.0. Cerebras serves GPT-OSS 120B near 3,000 tokens per second at $0.35 in and $0.75 out, plus Gemma 4 31B, on its wafer-scale chip. StepFun's pitch is cheap image and video understanding. Cerebras' pitch is raw speed.
Deployment and data location also differ. StepFun's first-party inference is China-hosted with thin Western distribution, though OpenRouter carries Step 3.7 Flash and the open weights run on vLLM and SGLang. Cerebras sells through a shared API, dedicated endpoints and partners including AWS Marketplace. StepFun trails frontier models on hard multimodal reasoning. Pick StepFun for cost-sensitive vision agents or cheap self-hosting. Pick Cerebras for voice and streaming text.
What Cerebras and StepFun do
Cerebras
Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.
Example models: GPT-OSS 120B, Gemma 4 31B
Full Cerebras profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose Cerebras or StepFun?
Cerebras
Choose Cerebras for
- Streaming text and voice at maximum speed
- Teams avoiding China-hosted inference
- Buying through AWS Marketplace
StepFun
Choose StepFun for
- Cheap image and video understanding in agents
- Self-hosting a small-active-parameter model under Apache 2.0
- 256K context multimodal work at $0.20 per million input
Cerebras vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Open (Apache 2.0) and API models |
| Flagship models | GPT-OSS 120B, Gemma 4 31B | Step 3.7 Flash, Step3 |
| Speed | ~3,000 tok/s on GPT-OSS 120B | ~128 tok/s on Step 3.7 Flash |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Unknown | Open weights to fine-tune |
| Deployment | Shared API, dedicated, partners | First-party API, OpenRouter |
| Long context | Unknown | 256K |
Frequently asked questions
What is the difference between Cerebras and StepFun?
StepFun ships cheap Apache 2.0 multimodal models from Shanghai. Cerebras serves a small set of text models at the fastest public speeds.
When should I choose Cerebras over StepFun?
Streaming text and voice at maximum speed; Teams avoiding China-hosted inference; Buying through AWS Marketplace.
When should I choose StepFun over Cerebras?
Cheap image and video understanding in agents; Self-hosting a small-active-parameter model under Apache 2.0; 256K context multimodal work at $0.20 per million input.
Is Cerebras or StepFun cheaper?
Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.