Google Vertex AI vs StepFun
A US hyperscaler platform against a Shanghai lab with cheap, Apache 2.0 multimodal models. Vertex wins on frontier quality and governance; StepFun on price and self-hosting.
By The Subconscious Team · Updated
Google Vertex AI vs StepFun: key differences
StepFun's main offer is efficient multimodal models. Step 3.7 Flash is a 198B mixture-of-experts vision-language model with only 11B active parameters, 256K context, selectable reasoning levels and structured outputs, priced at $0.20 in and $1.15 out on StepFun's API and released under Apache 2.0. Vertex AI covers multimodal work too, with Gemini 3.8 handling long-context and video tasks, plus Claude, Gemma and media models like Imagen and Veo inside a governed Google Cloud platform. The overlap is real, but the two sit at opposite ends on price and openness.
Quality and jurisdiction split them. StepFun trails frontier models on hard multimodal reasoning benchmarks, and its first-party inference is hosted in China with thin Western distribution and support. Its open weights run on vLLM and SGLang, though, so a team can self-host a small-active-parameter model cheaply wherever it likes. Vertex is the pick for hard multimodal reasoning and enterprise controls, with pricing that is harder to forecast. Cost-sensitive agents that need image and video understanding, and teams willing to host weights themselves, fit StepFun.
What Google Vertex AI and StepFun do
Google Vertex AI
Vertex AI is Google Cloud's enterprise AI platform. At Google Cloud Next on April 22, 2026, Google rebranded it the Gemini Enterprise Agent Platform with an agent-first structure, though the API endpoint and most docs still say Vertex. Model Garden offers 200+ models, including Google's Gemini 3.8 family, Anthropic's Claude models and open models like Gemma, alongside Imagen, Veo and Chirp for media and speech. Google's own TPUs sit underneath much of its first-party serving.
Example models: Gemini 3.8, Claude
Full Google Vertex AI profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose Google Vertex AI or StepFun?
Google Vertex AI
Choose Google Vertex AI for
- Hard multimodal reasoning and video on Gemini
- Enterprise controls and Google Cloud procurement
- A choice of closed and open models
StepFun
Choose StepFun for
- Cheap vision and video understanding in agents
- Self-hosting an Apache 2.0 model with 11B active
- Low per-token multimodal prices
Google Vertex AI vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Closed and open, 200+ models | Open (Apache 2.0) and API models |
| Flagship models | Gemini 3.8 Flash, Claude, Gemma | Step 3.7 Flash, Step3 |
| Speed | Flash tier built for low latency | ~128 tok/s on Step 3.7 Flash |
| Price | Gemini 3.8 Flash $0.75 in, $3.75 out | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Custom training on GPUs or TPUs | Open weights to fine-tune |
| Deployment | Managed on Google Cloud | First-party API, OpenRouter |
| Long context | 1M on Gemini 3.8 Flash | 256K |
Frequently asked questions
What is the difference between Google Vertex AI and StepFun?
A US hyperscaler platform against a Shanghai lab with cheap, Apache 2.0 multimodal models. Vertex wins on frontier quality and governance; StepFun on price and self-hosting.
When should I choose Google Vertex AI over StepFun?
Hard multimodal reasoning and video on Gemini; Enterprise controls and Google Cloud procurement; A choice of closed and open models.
When should I choose StepFun over Google Vertex AI?
Cheap vision and video understanding in agents; Self-hosting an Apache 2.0 model with 11B active; Low per-token multimodal prices.
Is Google Vertex AI or StepFun cheaper?
Google Vertex AI: Gemini 3.8 Flash $0.75 in, $3.75 out. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, Google Vertex AI or StepFun?
Google Vertex AI: 1M on Gemini 3.8 Flash. StepFun: 256K.
Related comparisons
Subconscious vs Google Vertex AI
OpenAI vs Google Vertex AI
Anthropic vs Google Vertex AI
Google Vertex AI vs Amazon Bedrock
Google Vertex AI vs Together AI
Google Vertex AI vs Fireworks AI
Subconscious vs StepFun
OpenAI vs StepFun
Anthropic vs StepFun
Amazon Bedrock vs StepFun
Together AI vs StepFun
Fireworks AI vs StepFun
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.