Subconscious vs StepFun
StepFun ships cheap multimodal models with a 256K window. Subconscious serves open models with a 5M+ effective context, built for agent traces far past that.
By The Subconscious Team · Updated
Subconscious vs StepFun: key differences
StepFun is a lab and Subconscious is an inference company, and they meet at efficient open models. StepFun's Step 3.7 Flash is a 198B mixture-of-experts vision-language model with only 11B active parameters, 256K context and Apache 2.0 weights, priced at $0.20 in and $1.15 out on its own API. It reads images and video cheaply. Its window ends at 256K, though, which a long coding agent can pass partway through a session. Subconscious serves GLM 5.3 and DeepSeek V4.1 Flash on a runtime that prunes the KV cache, delivers a 5M+ effective context window, and bills processed tokens rather than every token sent.
StepFun is the pick for vision and video understanding inside cost-sensitive agents, or for a small-active-parameter model to self-host under a permissive license. Its trade-offs are a gap to frontier models on hard multimodal reasoning and China-hosted first-party inference with thin Western support. Subconscious records no prompts or inputs and offers dedicated and on-prem deployments that can run nearly any open model. For agents that run for an hour over a large codebase, its 2x faster task completion and processed-token billing matter more than a low list price.
What Subconscious and StepFun do
Subconscious
Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.
Example models: GLM 5.3, DeepSeek V4.1 Flash
Full Subconscious profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose Subconscious or StepFun?
Subconscious
Choose Subconscious for
- Agent traces that run past StepFun's 256K window
- Teams avoiding China-hosted first-party inference
- Long coding sessions billed on processed tokens
StepFun
Choose StepFun for
- Cheap image and video understanding inside agents
- Self-hosting a small-active-parameter model under Apache 2.0
- Multimodal work that fits within a 256K window
Subconscious vs StepFun at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Open (Apache 2.0) and API models |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | Step 3.7 Flash, Step3 |
| Speed | 2x faster task completion | ~128 tok/s on Step 3.7 Flash |
| Price | 50–80% lower cost; billed on processed tokens | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | Marathon post-trained variants | Open weights to fine-tune |
| Deployment | Managed API, dedicated, on-prem | First-party API, OpenRouter |
| Long context | 5M+ effective context | 256K |
Frequently asked questions
What is the difference between Subconscious and StepFun?
StepFun ships cheap multimodal models with a 256K window. Subconscious serves open models with a 5M+ effective context, built for agent traces far past that.
When should I choose Subconscious over StepFun?
Agent traces that run past StepFun's 256K window; Teams avoiding China-hosted first-party inference; Long coding sessions billed on processed tokens.
When should I choose StepFun over Subconscious?
Cheap image and video understanding inside agents; Self-hosting a small-active-parameter model under Apache 2.0; Multimodal work that fits within a 256K window.
Is Subconscious or StepFun cheaper?
Subconscious: 50–80% lower cost; billed on processed tokens. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, Subconscious or StepFun?
Subconscious: 5M+ effective context. StepFun: 256K.
Related comparisons
Run your longest agent traces on Subconscious
Point the OpenAI or Anthropic SDK, or the coding agent you already use, at Subconscious. Keep StepFun for the work it does best and send the long runs to us.