Relace vs StepFun
StepFun makes general multimodal models; Relace makes code-apply and search tools for coding agents. Step 3.7 Flash could drive an agent that Relace speeds up.
By The Subconscious Team · Updated
Relace vs StepFun: key differences
These are not substitutes. StepFun is a Shanghai lab whose Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context, tool use and structured outputs, priced at $0.20 in and $1.15 out and released under Apache 2.0. Relace does no general model serving. It sells small models for coding-agent chores: relace-apply-3 merges edits at about 10,000 tokens per second with a 128K cap, agentic search explores large codebases in parallel, and compaction runs at 50,000 tokens per second.
A budget coding agent could combine them, with Step 3.7 Flash planning and writing edits, possibly working from screenshots thanks to its image input, and Relace applying the edits and fetching relevant code. StepFun's larger 256K window covers inputs that exceed Relace's 128K limit. Deployment differs too. Relace offers hosted or self-hosted options, while StepFun's first-party inference is China-hosted with thin Western support, though its open weights run on vLLM and SGLang. StepFun also trails frontier models on hard multimodal reasoning.
What Relace and StepFun do
Relace
Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.
Example models: relace-apply-3, Relace agentic search
Full Relace profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose Relace or StepFun?
Relace vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Specialist models | Open (Apache 2.0) and API models |
| Flagship models | relace-apply-3, agentic search | Step 3.7 Flash, Step3 |
| Speed | ~10,000 tok/s apply | ~128 tok/s on Step 3.7 Flash |
| Price | 3x+ cheaper than full rewrites | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Unknown | Open weights to fine-tune |
| Deployment | Hosted API or self-hosted | First-party API, OpenRouter |
| Long context | 128K max | 256K |
Frequently asked questions
What is the difference between Relace and StepFun?
StepFun makes general multimodal models; Relace makes code-apply and search tools for coding agents. Step 3.7 Flash could drive an agent that Relace speeds up.
When should I choose Relace over StepFun?
Fast merges of edit snippets; Codebase search for PR review and CI; Self-hosted utility models.
When should I choose StepFun over Relace?
A cheap, self-hostable main model for agents; Reading screenshots and video inside agents; Contexts up to 256K tokens.
Is Relace or StepFun cheaper?
Relace: 3x+ cheaper than full rewrites. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, Relace or StepFun?
Relace: 128K max. StepFun: 256K.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.