vs

Subconscious vs Relace

Relace makes fast tool models for coding agents, and its apply model stops at 128K tokens. Subconscious serves the main agent model on traces that run far past that.

By The Subconscious Team · Updated

Subconscious vs Relace: key differences

Relace sells utilities, not a general model. Its relace-apply-3 merges lazy edit snippets into files at about 10,000 tokens per second, its agentic search explores large codebases in parallel, and a compaction model runs at 50,000 tokens per second. Relace argues that small specialized models beat frontier LLMs on these utility jobs. Apply has a hard limit, though: it returns an error past 128K tokens. Subconscious is built for the context beyond that. It serves the agent's main model, prunes the KV cache on long traces, bills processed tokens, and delivers a 5M+ effective context window.

In a coding stack they divide the work. Subconscious runs the long planning and editing loop across the whole repository, where processed-token billing keeps an hour of accumulated context affordable. Relace handles the fast, narrow calls around it: applying edits, searching the repo and compacting context. Relace's compaction runs as a separate model call, while Subconscious prunes cache state inside the runtime, so the two work at different layers. Relace's self-hosted option and Subconscious's on-prem deployment both suit enterprises that keep code in-house, and for files past Relace's 128K cap, the main model on Subconscious is the obvious fallback.

What Subconscious and Relace do

Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

Example models: GLM 5.3, DeepSeek V4.1 Flash

Full Subconscious profile

Relace

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

Example models: relace-apply-3, Relace agentic search

Full Relace profile

Should you choose Subconscious or Relace?

Subconscious

Choose Subconscious for

  • The main agent model on repos past Relace's 128K cap
  • Hour-long coding loops billed on processed tokens
  • On-prem long-context serving

Relace

Choose Relace for

  • Instant apply of lazy edits at about 10,000 tokens per second
  • Fast parallel search across large codebases
  • Self-hosted coding utilities for in-house code

Subconscious vs Relace at a glance

AttributeSubconsciousRelace
Model accessOpen weightsSpecialist models
Flagship modelsGLM 5.3, DeepSeek V4.1 Flashrelace-apply-3, agentic search
Speed2x faster task completion~10,000 tok/s apply
Price50–80% lower cost; billed on processed tokens3x+ cheaper than full rewrites
CustomizationMarathon post-trained variantsUnknown
DeploymentManaged API, dedicated, on-premHosted API or self-hosted
Long context5M+ effective context128K max

Frequently asked questions

What is the difference between Subconscious and Relace?

Relace makes fast tool models for coding agents, and its apply model stops at 128K tokens. Subconscious serves the main agent model on traces that run far past that.

When should I choose Subconscious over Relace?

The main agent model on repos past Relace's 128K cap; Hour-long coding loops billed on processed tokens; On-prem long-context serving.

When should I choose Relace over Subconscious?

Instant apply of lazy edits at about 10,000 tokens per second; Fast parallel search across large codebases; Self-hosted coding utilities for in-house code.

Is Subconscious or Relace cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Relace: 3x+ cheaper than full rewrites. The cheaper choice depends on the model and workload.

Which has more context, Subconscious or Relace?

Subconscious: 5M+ effective context. Relace: 128K max.

Related comparisons

Run your longest agent traces on Subconscious

Point the OpenAI or Anthropic SDK, or the coding agent you already use, at Subconscious. Keep Relace for the work it does best and send the long runs to us.