We raised $5.1M for long-running agents.
vs

Mistral AI vs StepFun

Two labs shipping open weights with 256K context. StepFun's Step 3.7 Flash is a cheap vision-language MoE; Mistral has wider cloud reach and EU or US hosting.

By The Subconscious Team · Updated

Mistral AI vs StepFun: key differences

Both labs release open weights under permissive licenses and run their own APIs. StepFun's Step 3.7 Flash, from May 2026, is a 198B MoE vision-language model with 11B active, 256K context, selectable reasoning levels and tool use, under Apache 2.0 at $0.20 in and $1.15 out per million tokens and about 128 tokens per second. Mistral Small 4, a 119B MoE with 6.5B active, costs $0.15 in and $0.60 out, and Large 3, a 675B MoE also under Apache 2.0, costs $0.50 in and $1.50 out. For harder coding, Mistral has Medium 3.5 at $1.50 in and $7.50 out. StepFun's edge is image and video understanding in a small-active model.

Distribution is the bigger gap. StepFun's first-party inference is China-hosted, with thin Western distribution and support, and outside its own API it reaches developers mainly through OpenRouter. Mistral serves from EU or US regions, sells through Azure, Bedrock, Vertex AI, Snowflake and watsonx, and offers a Priority Tier with uptime SLAs. StepFun also trails frontier models on hard multimodal reasoning benchmarks. On customization, StepFun's open weights run on vLLM and SGLang and can be fine-tuned freely. Mistral's weights allow the same self-hosted path, but its managed fine-tuning API is deprecated in favor of enterprise Forge, and it retires models quickly, which forces regular migrations.

What Mistral AI and StepFun do

Mistral AI

Mistral AI is a Paris lab that sells its models through La Plateforme, its own API, and releases most of them as open weights. It consolidated the lineup in 2026. Mistral Medium 3.5, released April 28, is a dense 128B model that merges instruction following, reasoning and coding into one set of weights, and it replaced both Devstral 2 and the Magistral reasoning models. It costs $1.50 in and $7.50 out per million tokens and scores 77.6% on SWE-Bench Verified by Mistral's count. Mistral Small 4, a 119B mixture-of-experts model with 6.5B active, costs $0.15 in and $0.60 out. Mistral Large 3, a 675B MoE under Apache 2.0, runs $0.50 in and $1.50 out. All three carry a 256K context window.

Example models: Mistral Medium 3.5, Mistral Small 4

Full Mistral AI profile

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Should you choose Mistral AI or StepFun?

Mistral AI

Choose Mistral AI for

  • Western enterprises needing EU or US hosting
  • Agentic coding on Medium 3.5
  • Buying through hyperscaler marketplaces

StepFun

Choose StepFun for

  • Cheap vision and video understanding in agents
  • Self-hosting a small-active multimodal model
  • Selectable reasoning levels on one model

Mistral AI vs StepFun at a glance

AttributeMistral AIStepFun
Model accessOpen weights, plus closed CodestralOpen (Apache 2.0) and API models
Flagship modelsMistral Medium 3.5, Small 4, Large 3Step 3.7 Flash, Step3
SpeedUnknown~128 tok/s on Step 3.7 Flash
Price$0.15–$1.50 in, $0.60–$7.50 out per 1M$0.20 in, $1.15 out (Step 3.7 Flash)
CustomizationForge (enterprise); fine-tuning API deprecatedOpen weights to fine-tune
DeploymentAPI, Azure, Bedrock, Vertex, self-hostFirst-party API, OpenRouter
Long context256K256K

Frequently asked questions

What is the difference between Mistral AI and StepFun?

Two labs shipping open weights with 256K context. StepFun's Step 3.7 Flash is a cheap vision-language MoE; Mistral has wider cloud reach and EU or US hosting.

When should I choose Mistral AI over StepFun?

Western enterprises needing EU or US hosting; Agentic coding on Medium 3.5; Buying through hyperscaler marketplaces.

When should I choose StepFun over Mistral AI?

Cheap vision and video understanding in agents; Self-hosting a small-active multimodal model; Selectable reasoning levels on one model.

Is Mistral AI or StepFun cheaper?

Mistral AI: $0.15–$1.50 in, $0.60–$7.50 out per 1M. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

Which has more context, Mistral AI or StepFun?

Mistral AI: 256K. StepFun: 256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.