vs

Google Vertex AI vs StepFun

A US hyperscaler platform against a Shanghai lab with cheap, Apache 2.0 multimodal models. Vertex wins on frontier quality and governance; StepFun on price and self-hosting.

By The Subconscious Team · Updated

Google Vertex AI vs StepFun: key differences

StepFun's main offer is efficient multimodal models. Step 3.7 Flash is a 198B mixture-of-experts vision-language model with only 11B active parameters, 256K context, selectable reasoning levels and structured outputs, priced at $0.20 in and $1.15 out on StepFun's API and released under Apache 2.0. Vertex AI covers multimodal work too, with Gemini 3.8 handling long-context and video tasks, plus Claude, Gemma and media models like Imagen and Veo inside a governed Google Cloud platform. The overlap is real, but the two sit at opposite ends on price and openness.

Quality and jurisdiction split them. StepFun trails frontier models on hard multimodal reasoning benchmarks, and its first-party inference is hosted in China with thin Western distribution and support. Its open weights run on vLLM and SGLang, though, so a team can self-host a small-active-parameter model cheaply wherever it likes. Vertex is the pick for hard multimodal reasoning and enterprise controls, with pricing that is harder to forecast. Cost-sensitive agents that need image and video understanding, and teams willing to host weights themselves, fit StepFun.

What Google Vertex AI and StepFun do

Google Vertex AI

Vertex AI is Google Cloud's enterprise AI platform. At Google Cloud Next on April 22, 2026, Google rebranded it the Gemini Enterprise Agent Platform with an agent-first structure, though the API endpoint and most docs still say Vertex. Model Garden offers 200+ models, including Google's Gemini 3.8 family, Anthropic's Claude models and open models like Gemma, alongside Imagen, Veo and Chirp for media and speech. Google's own TPUs sit underneath much of its first-party serving.

Example models: Gemini 3.8, Claude

Full Google Vertex AI profile

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Should you choose Google Vertex AI or StepFun?

Google Vertex AI

Choose Google Vertex AI for

  • Hard multimodal reasoning and video on Gemini
  • Enterprise controls and Google Cloud procurement
  • A choice of closed and open models

StepFun

Choose StepFun for

  • Cheap vision and video understanding in agents
  • Self-hosting an Apache 2.0 model with 11B active
  • Low per-token multimodal prices

Google Vertex AI vs StepFun at a glance

AttributeGoogle Vertex AIStepFun
Model accessClosed and open, 200+ modelsOpen (Apache 2.0) and API models
Flagship modelsGemini 3.8 Flash, Claude, GemmaStep 3.7 Flash, Step3
SpeedFlash tier built for low latency~128 tok/s on Step 3.7 Flash
PriceGemini 3.8 Flash $0.75 in, $3.75 out$0.20 in, $1.15 out (Step 3.7 Flash)
CustomizationCustom training on GPUs or TPUsOpen weights to fine-tune
DeploymentManaged on Google CloudFirst-party API, OpenRouter
Long context1M on Gemini 3.8 Flash256K

Frequently asked questions

What is the difference between Google Vertex AI and StepFun?

A US hyperscaler platform against a Shanghai lab with cheap, Apache 2.0 multimodal models. Vertex wins on frontier quality and governance; StepFun on price and self-hosting.

When should I choose Google Vertex AI over StepFun?

Hard multimodal reasoning and video on Gemini; Enterprise controls and Google Cloud procurement; A choice of closed and open models.

When should I choose StepFun over Google Vertex AI?

Cheap vision and video understanding in agents; Self-hosting an Apache 2.0 model with 11B active; Low per-token multimodal prices.

Is Google Vertex AI or StepFun cheaper?

Google Vertex AI: Gemini 3.8 Flash $0.75 in, $3.75 out. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

Which has more context, Google Vertex AI or StepFun?

Google Vertex AI: 1M on Gemini 3.8 Flash. StepFun: 256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.