vs

GMI Cloud vs StepFun

StepFun sells its own low-cost multimodal models from China; GMI Cloud hosts many providers' models on owned hardware across the US and APAC.

By The Subconscious Team · Updated

GMI Cloud vs StepFun: key differences

StepFun is a lab. Its Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and Apache 2.0 weights, at $0.20 in and $1.15 out on StepFun's API. GMI Cloud is a host with its own GPUs and a catalog of 100+ models from providers like Google Veo, Kling, MiniMax and ElevenLabs, with entry text pricing such as GLM-4.7-Flash at $0.07 in and $0.40 out. StepFun also ships open weights that run on vLLM and SGLang, while GMI's value is the hardware and the multi-provider catalog.

Location and breadth separate them. StepFun's first-party inference is China-hosted with thin Western distribution. GMI runs in-country data centers in Taiwan, Thailand and Malaysia plus two US sites, which helps APAC companies that need data kept in a specific country. StepFun is the stronger pick for cheap image and video understanding straight from the model's makers. GMI suits apps that need video, image and audio generation alongside an LLM, or reserved GPU capacity to self-host open weights.

What GMI Cloud and StepFun do

GMI Cloud

GMI Cloud is a vertically integrated GPU cloud and inference platform that owns its NVIDIA hardware. It runs Tier-4 data centers in Silicon Valley, Colorado, Taiwan, Thailand and Malaysia, and as an NVIDIA Cloud Partner it gets priority access to H100, H200 and B200 supply. The company pivoted from crypto mining into AI, which gave it experience standing up high-density power and cooling fast. An $82M Series A came from Headline, Wistron and Thai energy group Banpu.

Example models: GLM-4.7-Flash, Google Veo

Full GMI Cloud profile

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Should you choose GMI Cloud or StepFun?

GMI Cloud

Choose GMI Cloud for

  • In-country APAC inference outside mainland China
  • Generating video, image and audio next to an LLM
  • Reserved H100 or H200 capacity for self-hosting

StepFun

Choose StepFun for

  • Low-cost vision and video understanding with 256K context
  • Apache 2.0 weights for cheap self-hosting
  • Selectable reasoning levels on a small-active model

GMI Cloud vs StepFun at a glance

AttributeGMI CloudStepFun
Model accessOpen and third-party modelsOpen (Apache 2.0) and API models
Flagship modelsGLM-4.7-Flash, Google VeoStep 3.7 Flash, Step3
SpeedNear bare-metal performance~128 tok/s on Step 3.7 Flash
Price$0.07 in, $0.40 out (GLM-4.7-Flash)$0.20 in, $1.15 out (Step 3.7 Flash)
CustomizationUnknownOpen weights to fine-tune
DeploymentShared, autoscaling, reserved GPUsFirst-party API, OpenRouter
Long contextVaries by model256K

Frequently asked questions

What is the difference between GMI Cloud and StepFun?

StepFun sells its own low-cost multimodal models from China; GMI Cloud hosts many providers' models on owned hardware across the US and APAC.

When should I choose GMI Cloud over StepFun?

In-country APAC inference outside mainland China; Generating video, image and audio next to an LLM; Reserved H100 or H200 capacity for self-hosting.

When should I choose StepFun over GMI Cloud?

Low-cost vision and video understanding with 256K context; Apache 2.0 weights for cheap self-hosting; Selectable reasoning levels on a small-active model.

Is GMI Cloud or StepFun cheaper?

GMI Cloud: $0.07 in, $0.40 out (GLM-4.7-Flash). StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

Which has more context, GMI Cloud or StepFun?

GMI Cloud: Varies by model. StepFun: 256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.