StepFun vs StreamLake
Two Chinese AI providers with different specialties. StepFun ships open multimodal models; StreamLake sells Kuaishou's proprietary coding model and cloud capacity.
By The Subconscious Team · Updated
StepFun vs StreamLake: key differences
Both providers serve first-party inference from China, which sets the same baseline caution for US and EU buyers. Past that, they diverge. StepFun is a Shanghai lab with efficient multimodal models, led by Step 3.7 Flash, a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and Apache 2.0 weights, at $0.20 in and $1.15 out. StreamLake is Kuaishou's AI cloud, and its headline model is KAT-Coder-Pro V2.5, a proprietary agentic coding model that Kuaishou trained with large-scale agentic reinforcement learning for repository-level work.
The choice follows the job. For coding agents, StreamLake offers a KwaiKAT Coding Plan and a Claude-protocol proxy that drops into Claude Code or OpenClaw. For vision, video and speech work, StepFun's models are built for it, and its open weights run on vLLM and SGLang wherever a team wants. StreamLake adds bare-metal compute for Chinese internet businesses, while StepFun's models reach Western users through OpenRouter. StreamLake's pricing and docs lead with China and yuan, and StepFun trails frontier models on hard multimodal reasoning.
What StepFun and StreamLake do
StepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileStreamLake
StreamLake is the AI cloud brand of Kuaishou, the Chinese short-video company behind the Kling video models. It sells model-as-a-service inference and bare-metal compute to internet businesses, drawing on the infrastructure Kuaishou built to serve video at massive scale. Its developer site offers APIs, SDKs and integration guides aimed at taking teams from testing to production.
Example models: KAT-Coder-Pro V2.5, KAT-Coder-Air
Full StreamLake profileShould you choose StepFun or StreamLake?
StepFun
Choose StepFun for
- Vision, video and speech understanding at low cost
- Open weights to self-host outside China
- Access through OpenRouter
StreamLake
Choose StreamLake for
- Agentic coding on KAT-Coder in Claude Code
- A flat KwaiKAT Coding Plan
- Bare-metal capacity for businesses in China
StepFun vs StreamLake at a glance
| Attribute | ||
|---|---|---|
| Model access | Open (Apache 2.0) and API models | Proprietary coding models |
| Flagship models | Step 3.7 Flash, Step3 | KAT-Coder-Pro V2.5, KAT-Coder-Air |
| Speed | ~128 tok/s on Step 3.7 Flash | Unknown |
| Price | $0.20 in, $1.15 out (Step 3.7 Flash) | Per token or KwaiKAT Coding Plan |
| Customization | Open weights to fine-tune | Unknown |
| Deployment | First-party API, OpenRouter | MaaS API, bare metal |
| Long context | 256K | Unknown |
Frequently asked questions
What is the difference between StepFun and StreamLake?
Two Chinese AI providers with different specialties. StepFun ships open multimodal models; StreamLake sells Kuaishou's proprietary coding model and cloud capacity.
When should I choose StepFun over StreamLake?
Vision, video and speech understanding at low cost; Open weights to self-host outside China; Access through OpenRouter.
When should I choose StreamLake over StepFun?
Agentic coding on KAT-Coder in Claude Code; A flat KwaiKAT Coding Plan; Bare-metal capacity for businesses in China.
Is StepFun or StreamLake cheaper?
StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). StreamLake: Per token or KwaiKAT Coding Plan. The cheaper choice depends on the model and workload.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.