# Subconscious vs Relace

> Relace makes fast tool models for coding agents, and its apply model stops at 128K tokens. Subconscious serves the main agent model on traces that run far past that.

Canonical: https://www.subconscious.dev/compare/subconscious-vs-relace · By The Subconscious Team · Updated September 30, 2026

## How they compare

Relace sells utilities, not a general model. Its relace-apply-3 merges lazy edit snippets into files at about 10,000 tokens per second, its agentic search explores large codebases in parallel, and a compaction model runs at 50,000 tokens per second. Relace argues that small specialized models beat frontier LLMs on these utility jobs. Apply has a hard limit, though: it returns an error past 128K tokens. Subconscious is built for the context beyond that. It serves the agent's main model, prunes the KV cache on long traces, bills processed tokens, and delivers a 5M+ effective context window.

In a coding stack they divide the work. Subconscious runs the long planning and editing loop across the whole repository, where processed-token billing keeps an hour of accumulated context affordable. Relace handles the fast, narrow calls around it: applying edits, searching the repo and compacting context. Relace's compaction runs as a separate model call, while Subconscious prunes cache state inside the runtime, so the two work at different layers. Relace's self-hosted option and Subconscious's on-prem deployment both suit enterprises that keep code in-house, and for files past Relace's 128K cap, the main model on Subconscious is the obvious fallback.

## What each one does

### Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### Relace

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

## Which is best, and when

### Choose Subconscious for

- The main agent model on repos past Relace's 128K cap
- Hour-long coding loops billed on processed tokens
- On-prem long-context serving

### Choose Relace for

- Instant apply of lazy edits at about 10,000 tokens per second
- Fast parallel search across large codebases
- Self-hosted coding utilities for in-house code

## At a glance

| Attribute | Subconscious | Relace |
|---|---|---|
| Model access | Open weights | Specialist models |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | relace-apply-3, agentic search |
| Speed | 2x faster task completion | ~10,000 tok/s apply |
| Price | 50–80% lower cost; billed on processed tokens | 3x+ cheaper than full rewrites |
| Customization | Marathon post-trained variants | - |
| Deployment | Managed API, dedicated, on-prem | Hosted API or self-hosted |
| Long context | 5M+ effective context | 128K max |

## FAQ

### What is the difference between Subconscious and Relace?

Relace makes fast tool models for coding agents, and its apply model stops at 128K tokens. Subconscious serves the main agent model on traces that run far past that.

### When should I choose Subconscious over Relace?

The main agent model on repos past Relace's 128K cap; Hour-long coding loops billed on processed tokens; On-prem long-context serving.

### When should I choose Relace over Subconscious?

Instant apply of lazy edits at about 10,000 tokens per second; Fast parallel search across large codebases; Self-hosted coding utilities for in-house code.

### Is Subconscious or Relace cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Relace: 3x+ cheaper than full rewrites. The cheaper choice depends on the model and workload.

### Which has more context, Subconscious or Relace?

Subconscious: 5M+ effective context. Relace: 128K max.

## Try Subconscious

Subconscious speaks the OpenAI and Anthropic API formats. Base URL: https://api.subconscious.dev/v1. Docs: https://docs.subconscious.dev. Get an API key: https://platform.subconscious.dev/signin. Agent guide: https://www.subconscious.dev/agents.md.

Full profiles: [Subconscious](https://www.subconscious.dev/providers/subconscious.md), [Relace](https://www.subconscious.dev/providers/relace.md).
