vs

Morph vs Relace

The closest head-to-head in coding-agent tooling. Both sell fast apply models and repo search; Morph edges on speed and breadth, Relace on self-hosting and compaction speed.

By The Subconscious Team · Updated

Morph vs Relace: key differences

These two really are substitutes. Both sell small models that merge a frontier model's partial edit into a full file, and both have grown into toolkits for coding agents. Morph's Fast Apply runs at 10,500+ tokens per second with up to 98% accuracy on a 7B model trained only on code merging. Relace's relace-apply-3 runs at about 10,000 tokens per second with 128K tokens of input and output. Morph says its approach cuts token usage about 40% against full rewrites. Relace says its apply is over 3x faster and cheaper than having the big model rewrite the file.

The toolkits differ at the edges. Morph adds WarpGrep for agentic repo search, Compact for context compression, Reflex for classification, fine-tuning and general chat endpoints. Relace adds parallel agentic search, a compaction model at 50,000 tokens per second, and source control with retrieval built in. Deployment is the clearest split: Relace offers self-hosting with guided onboarding for enterprises that keep code in-house, while Morph is an OpenAI-compatible hosted API. Relace errors past 128K tokens. Morph's 2 to 4% merge error rate calls for tests or linting.

What Morph and Relace do

Morph

Morph builds small, very fast specialist models that sit beside a big coding model inside an agent. Its flagship is Fast Apply. The frontier model writes only the changed lines with // ... existing code ... markers, and Morph merges them into the full file at 10,500+ tokens per second with up to 98% accuracy. It is the same idea behind Cursor's instant apply, offered as an OpenAI-compatible API.

Example models: morph-v3-fast, morph-v3-large

Full Morph profile

Relace

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

Example models: relace-apply-3, Relace agentic search

Full Relace profile

Should you choose Morph or Relace?

Morph

Choose Morph for

  • The fastest published apply speed, at 10,500+ tokens per second
  • One vendor for apply, search, classification and fine-tuning
  • CI and sandbox pipelines editing code at volume

Relace

Choose Relace for

  • Self-hosted apply and search for in-house code
  • Very fast context compaction at 50,000 tokens per second
  • App builders that need source control with retrieval

Morph vs Relace at a glance

AttributeMorphRelace
Model accessSpecialist modelsSpecialist models
Flagship modelsmorph-v3-fast, morph-v3-largerelace-apply-3, agentic search
Speed10,500+ tok/s Fast Apply~10,000 tok/s apply
Price~40% fewer tokens than full rewrites3x+ cheaper than full rewrites
CustomizationFine-tuning offeredUnknown
DeploymentOpenAI-compatible APIHosted API or self-hosted
Long contextUnknown128K max

Frequently asked questions

What is the difference between Morph and Relace?

The closest head-to-head in coding-agent tooling. Both sell fast apply models and repo search; Morph edges on speed and breadth, Relace on self-hosting and compaction speed.

When should I choose Morph over Relace?

The fastest published apply speed, at 10,500+ tokens per second; One vendor for apply, search, classification and fine-tuning; CI and sandbox pipelines editing code at volume.

When should I choose Relace over Morph?

Self-hosted apply and search for in-house code; Very fast context compaction at 50,000 tokens per second; App builders that need source control with retrieval.

Is Morph or Relace cheaper?

Morph: ~40% fewer tokens than full rewrites. Relace: 3x+ cheaper than full rewrites. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.