Subconscious vs Relace
Relace makes fast tool models for coding agents, and its apply model stops at 128K tokens. Subconscious serves the main agent model on traces that run far past that.
By The Subconscious Team · Updated
Subconscious vs Relace: key differences
Relace sells utilities, not a general model. Its relace-apply-3 merges lazy edit snippets into files at about 10,000 tokens per second, its agentic search explores large codebases in parallel, and a compaction model runs at 50,000 tokens per second. Relace argues that small specialized models beat frontier LLMs on these utility jobs. Apply has a hard limit, though: it returns an error past 128K tokens. Subconscious is built for the context beyond that. It serves the agent's main model, prunes the KV cache on long traces, bills processed tokens, and delivers a 5M+ effective context window.
In a coding stack they divide the work. Subconscious runs the long planning and editing loop across the whole repository, where processed-token billing keeps an hour of accumulated context affordable. Relace handles the fast, narrow calls around it: applying edits, searching the repo and compacting context. Relace's compaction runs as a separate model call, while Subconscious prunes cache state inside the runtime, so the two work at different layers. Relace's self-hosted option and Subconscious's on-prem deployment both suit enterprises that keep code in-house, and for files past Relace's 128K cap, the main model on Subconscious is the obvious fallback.
What Subconscious and Relace do
Subconscious
Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.
Example models: GLM 5.3, DeepSeek V4.1 Flash
Full Subconscious profileRelace
Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.
Example models: relace-apply-3, Relace agentic search
Full Relace profileShould you choose Subconscious or Relace?
Subconscious
Choose Subconscious for
- The main agent model on repos past Relace's 128K cap
- Hour-long coding loops billed on processed tokens
- On-prem long-context serving
Relace
Choose Relace for
- Instant apply of lazy edits at about 10,000 tokens per second
- Fast parallel search across large codebases
- Self-hosted coding utilities for in-house code
Subconscious vs Relace at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Specialist models |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | relace-apply-3, agentic search |
| Speed | 2x faster task completion | ~10,000 tok/s apply |
| Price | 50–80% lower cost; billed on processed tokens | 3x+ cheaper than full rewrites |
| Customization | Marathon post-trained variants | Unknown |
| Deployment | Managed API, dedicated, on-prem | Hosted API or self-hosted |
| Long context | 5M+ effective context | 128K max |
Frequently asked questions
What is the difference between Subconscious and Relace?
Relace makes fast tool models for coding agents, and its apply model stops at 128K tokens. Subconscious serves the main agent model on traces that run far past that.
When should I choose Subconscious over Relace?
The main agent model on repos past Relace's 128K cap; Hour-long coding loops billed on processed tokens; On-prem long-context serving.
When should I choose Relace over Subconscious?
Instant apply of lazy edits at about 10,000 tokens per second; Fast parallel search across large codebases; Self-hosted coding utilities for in-house code.
Is Subconscious or Relace cheaper?
Subconscious: 50–80% lower cost; billed on processed tokens. Relace: 3x+ cheaper than full rewrites. The cheaper choice depends on the model and workload.
Which has more context, Subconscious or Relace?
Subconscious: 5M+ effective context. Relace: 128K max.
Related comparisons
Run your longest agent traces on Subconscious
Point the OpenAI or Anthropic SDK, or the coding agent you already use, at Subconscious. Keep Relace for the work it does best and send the long runs to us.