vs

Baseten vs StepFun

StepFun is a Shanghai lab with cheap, efficient multimodal models, many under Apache 2.0. Baseten is a host that could serve those weights outside China.

By The Subconscious Team · Updated

Baseten vs StepFun: key differences

StepFun builds models; Baseten serves them. StepFun's workhorse, Step 3.7 Flash, is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context and an Apache 2.0 license, priced at $0.20 in and $1.15 out on StepFun's own API. That is cheap for a model that reads images and video. The weak spots are distribution and location: StepFun's first-party inference is China-hosted, and Western support is thin. Step 3.7 Flash is not in Baseten's 13-model list, but its open weights could run on a Baseten dedicated deployment through Truss.

That route matters mainly for companies that want StepFun's efficiency without sending data to China. Baseten brings HIPAA, data residency, a 99.99% SLA and the lowest measured time to first token, at the cost of per-minute GPU billing (an H100 at about $6.50 an hour) instead of StepFun's per-token rates. StepFun's own benchmarks trail frontier models on hard multimodal reasoning. For a cost-sensitive vision agent with no residency concerns, calling StepFun directly is simpler.

What Baseten and StepFun do

Baseten

Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.

Example models: GLM 5.2, gpt-oss 120B

Full Baseten profile

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Should you choose Baseten or StepFun?

Baseten

Choose Baseten for

  • Serving Apache 2.0 Step weights with data residency
  • Curated text models like DeepSeek V4 and GLM 5.2
  • Latency-bound production under a 99.99% SLA

StepFun

Choose StepFun for

  • Cheap image and video understanding at 256K context
  • Pay-per-token access with no GPU setup
  • Small-active-parameter models for self-hosting

Baseten vs StepFun at a glance

AttributeBasetenStepFun
Model accessOpen weights, 13 curatedOpen (Apache 2.0) and API models
Flagship modelsGLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120BStep 3.7 Flash, Step3
Speed0.49s TTFT, lowest measured~128 tok/s on Step 3.7 Flash
PriceH100 about $6.50/hr dedicated$0.20 in, $1.15 out (Step 3.7 Flash)
CustomizationDeploy any model with TrussOpen weights to fine-tune
DeploymentModel APIs, dedicated, self-hostFirst-party API, OpenRouter
Long contextVaries by model256K

Frequently asked questions

What is the difference between Baseten and StepFun?

StepFun is a Shanghai lab with cheap, efficient multimodal models, many under Apache 2.0. Baseten is a host that could serve those weights outside China.

When should I choose Baseten over StepFun?

Serving Apache 2.0 Step weights with data residency; Curated text models like DeepSeek V4 and GLM 5.2; Latency-bound production under a 99.99% SLA.

When should I choose StepFun over Baseten?

Cheap image and video understanding at 256K context; Pay-per-token access with no GPU setup; Small-active-parameter models for self-hosting.

Is Baseten or StepFun cheaper?

Baseten: H100 about $6.50/hr dedicated. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

Which has more context, Baseten or StepFun?

Baseten: Varies by model. StepFun: 256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.