vs

Relace vs StepFun

StepFun makes general multimodal models; Relace makes code-apply and search tools for coding agents. Step 3.7 Flash could drive an agent that Relace speeds up.

By The Subconscious Team · Updated

Relace vs StepFun: key differences

These are not substitutes. StepFun is a Shanghai lab whose Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context, tool use and structured outputs, priced at $0.20 in and $1.15 out and released under Apache 2.0. Relace does no general model serving. It sells small models for coding-agent chores: relace-apply-3 merges edits at about 10,000 tokens per second with a 128K cap, agentic search explores large codebases in parallel, and compaction runs at 50,000 tokens per second.

A budget coding agent could combine them, with Step 3.7 Flash planning and writing edits, possibly working from screenshots thanks to its image input, and Relace applying the edits and fetching relevant code. StepFun's larger 256K window covers inputs that exceed Relace's 128K limit. Deployment differs too. Relace offers hosted or self-hosted options, while StepFun's first-party inference is China-hosted with thin Western support, though its open weights run on vLLM and SGLang. StepFun also trails frontier models on hard multimodal reasoning.

What Relace and StepFun do

Relace

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

Example models: relace-apply-3, Relace agentic search

Full Relace profile

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Should you choose Relace or StepFun?

Relace

Choose Relace for

  • Fast merges of edit snippets
  • Codebase search for PR review and CI
  • Self-hosted utility models

StepFun

Choose StepFun for

  • A cheap, self-hostable main model for agents
  • Reading screenshots and video inside agents
  • Contexts up to 256K tokens

Relace vs StepFun at a glance

AttributeRelaceStepFun
Model accessSpecialist modelsOpen (Apache 2.0) and API models
Flagship modelsrelace-apply-3, agentic searchStep 3.7 Flash, Step3
Speed~10,000 tok/s apply~128 tok/s on Step 3.7 Flash
Price3x+ cheaper than full rewrites$0.20 in, $1.15 out (Step 3.7 Flash)
CustomizationUnknownOpen weights to fine-tune
DeploymentHosted API or self-hostedFirst-party API, OpenRouter
Long context128K max256K

Frequently asked questions

What is the difference between Relace and StepFun?

StepFun makes general multimodal models; Relace makes code-apply and search tools for coding agents. Step 3.7 Flash could drive an agent that Relace speeds up.

When should I choose Relace over StepFun?

Fast merges of edit snippets; Codebase search for PR review and CI; Self-hosted utility models.

When should I choose StepFun over Relace?

A cheap, self-hostable main model for agents; Reading screenshots and video inside agents; Contexts up to 256K tokens.

Is Relace or StepFun cheaper?

Relace: 3x+ cheaper than full rewrites. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

Which has more context, Relace or StepFun?

Relace: 128K max. StepFun: 256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.