xAI vs StepFun
A closed US lab with live X data against a Shanghai lab with cheap, Apache 2.0 multimodal models. Grok for fresh data and reasoning; StepFun for low-cost vision.
By The Subconscious Team · Updated
xAI vs StepFun: key differences
StepFun's lead model, Step 3.7 Flash, is a 198B mixture-of-experts vision-language model with only 11B active parameters, 256K context and selectable reasoning levels, priced at $0.20 in and $1.15 out per million under an Apache 2.0 license. xAI's Grok 4.6 costs $2 in and $6 out with 500K context, and Grok 4.20 offers 1M at $1.25 in and $2.50 out. StepFun is far cheaper per token. xAI brings native X Search, a coding model called grok-build, and first-party image, video and audio APIs.
Quality and jurisdiction push in xAI's favor for some buyers. StepFun trails frontier models on hard multimodal reasoning benchmarks, and its first-party inference is China-hosted with thin Western distribution and support. Its open weights run on vLLM and SGLang, though, so teams can self-host anywhere. xAI is closed and its enterprise footprint is smaller than the biggest labs, and its price doubles past 200K prompt tokens. Cost-sensitive vision and video understanding fits StepFun. Real-time agents fit Grok.
What xAI and StepFun do
xAI
xAI sells the Grok models through its own API. Grok 4.6 is the current flagship and xAI tells developers to use it for everything outside audio, image and video, code included. It has a 500K context window and costs $2 in and $6 out per million tokens under 200K prompt tokens. Older Grok 4.20 and 4.3 models keep a 1M window at $1.25 in and $2.50 out, which is aggressive for that capability class.
Example models: Grok 4.6, Grok 4.20
Full xAI profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose xAI or StepFun?
xAI vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Closed | Open (Apache 2.0) and API models |
| Flagship models | Grok 4.6, Grok 4.20, grok-build | Step 3.7 Flash, Step3 |
| Speed | ~54 tok/s on Grok 4.6 | ~128 tok/s on Step 3.7 Flash |
| Price | $2 in, $6 out (Grok 4.6); 2x past 200K | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Unknown | Open weights to fine-tune |
| Deployment | First-party API | First-party API, OpenRouter |
| Long context | 500K (4.6), 1M (4.20, 4.3) | 256K |
Frequently asked questions
What is the difference between xAI and StepFun?
A closed US lab with live X data against a Shanghai lab with cheap, Apache 2.0 multimodal models. Grok for fresh data and reasoning; StepFun for low-cost vision.
When should I choose xAI over StepFun?
Real-time agents on X and web data; Contexts beyond 256K on Grok 4.20; Teams that avoid China-hosted inference.
When should I choose StepFun over xAI?
Low-cost image and video understanding; Self-hosting a model with 11B active parameters; Apache 2.0 weights for fine-tuning.
Is xAI or StepFun cheaper?
xAI: $2 in, $6 out (Grok 4.6); 2x past 200K. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, xAI or StepFun?
xAI: 500K (4.6), 1M (4.20, 4.3). StepFun: 256K.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.