GMI Cloud vs StepFun
StepFun sells its own low-cost multimodal models from China; GMI Cloud hosts many providers' models on owned hardware across the US and APAC.
By The Subconscious Team · Updated
GMI Cloud vs StepFun: key differences
StepFun is a lab. Its Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and Apache 2.0 weights, at $0.20 in and $1.15 out on StepFun's API. GMI Cloud is a host with its own GPUs and a catalog of 100+ models from providers like Google Veo, Kling, MiniMax and ElevenLabs, with entry text pricing such as GLM-4.7-Flash at $0.07 in and $0.40 out. StepFun also ships open weights that run on vLLM and SGLang, while GMI's value is the hardware and the multi-provider catalog.
Location and breadth separate them. StepFun's first-party inference is China-hosted with thin Western distribution. GMI runs in-country data centers in Taiwan, Thailand and Malaysia plus two US sites, which helps APAC companies that need data kept in a specific country. StepFun is the stronger pick for cheap image and video understanding straight from the model's makers. GMI suits apps that need video, image and audio generation alongside an LLM, or reserved GPU capacity to self-host open weights.
What GMI Cloud and StepFun do
GMI Cloud
GMI Cloud is a vertically integrated GPU cloud and inference platform that owns its NVIDIA hardware. It runs Tier-4 data centers in Silicon Valley, Colorado, Taiwan, Thailand and Malaysia, and as an NVIDIA Cloud Partner it gets priority access to H100, H200 and B200 supply. The company pivoted from crypto mining into AI, which gave it experience standing up high-density power and cooling fast. An $82M Series A came from Headline, Wistron and Thai energy group Banpu.
Example models: GLM-4.7-Flash, Google Veo
Full GMI Cloud profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose GMI Cloud or StepFun?
GMI Cloud
Choose GMI Cloud for
- In-country APAC inference outside mainland China
- Generating video, image and audio next to an LLM
- Reserved H100 or H200 capacity for self-hosting
StepFun
Choose StepFun for
- Low-cost vision and video understanding with 256K context
- Apache 2.0 weights for cheap self-hosting
- Selectable reasoning levels on a small-active model
GMI Cloud vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Open and third-party models | Open (Apache 2.0) and API models |
| Flagship models | GLM-4.7-Flash, Google Veo | Step 3.7 Flash, Step3 |
| Speed | Near bare-metal performance | ~128 tok/s on Step 3.7 Flash |
| Price | $0.07 in, $0.40 out (GLM-4.7-Flash) | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Unknown | Open weights to fine-tune |
| Deployment | Shared, autoscaling, reserved GPUs | First-party API, OpenRouter |
| Long context | Varies by model | 256K |
Frequently asked questions
What is the difference between GMI Cloud and StepFun?
StepFun sells its own low-cost multimodal models from China; GMI Cloud hosts many providers' models on owned hardware across the US and APAC.
When should I choose GMI Cloud over StepFun?
In-country APAC inference outside mainland China; Generating video, image and audio next to an LLM; Reserved H100 or H200 capacity for self-hosting.
When should I choose StepFun over GMI Cloud?
Low-cost vision and video understanding with 256K context; Apache 2.0 weights for cheap self-hosting; Selectable reasoning levels on a small-active model.
Is GMI Cloud or StepFun cheaper?
GMI Cloud: $0.07 in, $0.40 out (GLM-4.7-Flash). StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, GMI Cloud or StepFun?
GMI Cloud: Varies by model. StepFun: 256K.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.