Groq vs StepFun
StepFun is a Shanghai lab with cheap multimodal Step models under Apache 2.0. Groq is a speed host for a few text models on its own chip.
By The Subconscious Team · Updated
Groq vs StepFun: key differences
StepFun makes models and Groq hosts them, though Groq does not host StepFun's. Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active, 256K context and selectable reasoning levels, priced at $0.20 in and $1.15 out on StepFun's API under an Apache 2.0 license. It reads images and video. Groq's models are text only apart from Whisper, and its context caps around 131K. On multimodal input and context length, StepFun covers more. On raw speed and latency consistency, Groq's LPU is in another tier.
Location and support matter here. StepFun's first-party inference is hosted in China, and its Western distribution is thin, though OpenRouter carries its model. Groq's API is OpenAI-compatible, widely used and near the floor on small-model prices. StepFun's open weights can be fine-tuned and self-hosted on vLLM or SGLang, while Groq hosts no fine-tunes. Pick StepFun for cheap vision agents. Pick Groq for fast text.
What Groq and StepFun do
Groq
Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.
Example models: GPT-OSS 120B, Qwen 3.6 27B
Full Groq profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose Groq or StepFun?
Groq vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Open (Apache 2.0) and API models |
| Flagship models | GPT-OSS 120B, Qwen 3.6 27B | Step 3.7 Flash, Step3 |
| Speed | 500–1,000 tok/s | ~128 tok/s on Step 3.7 Flash |
| Price | Near the floor on small models | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | No fine-tuned model hosting | Open weights to fine-tune |
| Deployment | GroqCloud API | First-party API, OpenRouter |
| Long context | Around 131K max | 256K |
Frequently asked questions
What is the difference between Groq and StepFun?
StepFun is a Shanghai lab with cheap multimodal Step models under Apache 2.0. Groq is a speed host for a few text models on its own chip.
When should I choose Groq over StepFun?
Fast text generation for real-time apps; Voice pipelines with Whisper; Skipping China-hosted first-party inference.
When should I choose StepFun over Groq?
Cheap image and video understanding; Long 256K context on a multimodal model; Self-hosting Apache 2.0 weights.
Is Groq or StepFun cheaper?
Groq: Near the floor on small models. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, Groq or StepFun?
Groq: Around 131K max. StepFun: 256K.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.