# Relace

> Fast small models that act as tools for coding agents: apply, search and compaction.

Canonical: https://www.subconscious.dev/providers/relace · By The Subconscious Team · Updated September 30, 2026

- Example models: relace-apply-3, Relace agentic search
- Website: https://www.relace.ai

## Overview

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

Relace has grown into a broader toolkit for agent builders. Its agentic search explores large codebases in parallel and answers in seconds, a context compaction model runs at 50,000 tokens per second, and it offers source control with codebase retrieval built in. The company argues that specialized small models beat frontier LLMs on these utility tasks while cutting agent cost. Teams can use the hosted API or deploy self-hosted with guided onboarding.

## Upsides

- Very fast merges and retrieval that take work off expensive frontier models.
- Self-hosted deployment for enterprises that keep code in-house.

## Core use cases

- App builders and coding agents that apply AI edits to user codebases.
- PR review, automated fixes and CI pipelines that search large repos.

## Downsides

- Point solution for coding workflows with no general-purpose model serving.
- Returns an error past 128K tokens, so very large files need a fallback model.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Specialist models |
| Flagship models | relace-apply-3, agentic search |
| Speed | ~10,000 tok/s apply |
| Price | 3x+ cheaper than full rewrites |
| Customization | - |
| Deployment | Hosted API or self-hosted |
| Long context | 128K max |

## FAQ

### What is Relace?

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

### What is Relace best for?

App builders and coding agents that apply AI edits to user codebases; PR review, automated fixes and CI pipelines that search large repos.

### How much does Relace cost?

Relace pricing at a glance: 3x+ cheaper than full rewrites. Rates change often, so check Relace's pricing page before committing.

### How much context does Relace support?

Relace's long-context support: 128K max.

### What are the downsides of Relace?

Point solution for coding workflows with no general-purpose model serving; Returns an error past 128K tokens, so very large files need a fallback model.

### What are the best alternatives to Relace?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Relace on this site.

## Comparisons

- [Subconscious vs Relace](https://www.subconscious.dev/compare/subconscious-vs-relace.md)
- [OpenAI vs Relace](https://www.subconscious.dev/compare/openai-vs-relace.md)
- [Anthropic vs Relace](https://www.subconscious.dev/compare/anthropic-vs-relace.md)
- [Google Vertex AI vs Relace](https://www.subconscious.dev/compare/google-vertex-vs-relace.md)
- [Amazon Bedrock vs Relace](https://www.subconscious.dev/compare/aws-bedrock-vs-relace.md)
- [Together AI vs Relace](https://www.subconscious.dev/compare/together-ai-vs-relace.md)
- [Fireworks AI vs Relace](https://www.subconscious.dev/compare/fireworks-vs-relace.md)
- [Baseten vs Relace](https://www.subconscious.dev/compare/baseten-vs-relace.md)
- [Groq vs Relace](https://www.subconscious.dev/compare/groq-vs-relace.md)
- [Cerebras vs Relace](https://www.subconscious.dev/compare/cerebras-vs-relace.md)
- [DeepInfra vs Relace](https://www.subconscious.dev/compare/deepinfra-vs-relace.md)
- [Modal vs Relace](https://www.subconscious.dev/compare/modal-vs-relace.md)
- [xAI vs Relace](https://www.subconscious.dev/compare/xai-vs-relace.md)
- [DeepSeek vs Relace](https://www.subconscious.dev/compare/deepseek-vs-relace.md)
- [Moonshot AI vs Relace](https://www.subconscious.dev/compare/moonshot-ai-vs-relace.md)
- [Z.ai vs Relace](https://www.subconscious.dev/compare/z-ai-vs-relace.md)
- [Alibaba Cloud vs Relace](https://www.subconscious.dev/compare/alibaba-cloud-vs-relace.md)
- [Meta vs Relace](https://www.subconscious.dev/compare/meta-vs-relace.md)
- [SambaNova vs Relace](https://www.subconscious.dev/compare/sambanova-vs-relace.md)
- [Nebius vs Relace](https://www.subconscious.dev/compare/nebius-vs-relace.md)
- [fal vs Relace](https://www.subconscious.dev/compare/fal-vs-relace.md)
- [Novita AI vs Relace](https://www.subconscious.dev/compare/novita-ai-vs-relace.md)
- [Parasail vs Relace](https://www.subconscious.dev/compare/parasail-vs-relace.md)
- [Inference.net vs Relace](https://www.subconscious.dev/compare/inference-net-vs-relace.md)
- [GMI Cloud vs Relace](https://www.subconscious.dev/compare/gmi-cloud-vs-relace.md)
- [Sail Research vs Relace](https://www.subconscious.dev/compare/sail-research-vs-relace.md)
- [Morph vs Relace](https://www.subconscious.dev/compare/morph-vs-relace.md)
- [Relace vs TypeSafe AI](https://www.subconscious.dev/compare/relace-vs-typesafe-ai.md)
- [Relace vs StepFun](https://www.subconscious.dev/compare/relace-vs-stepfun.md)
- [Relace vs Runware](https://www.subconscious.dev/compare/relace-vs-runware.md)
- [Relace vs StreamLake](https://www.subconscious.dev/compare/relace-vs-streamlake.md)
- [Relace vs Wafer](https://www.subconscious.dev/compare/relace-vs-wafer.md)
- [Relace vs RunInfra](https://www.subconscious.dev/compare/relace-vs-runinfra.md)
- [Relace vs Particle.AI](https://www.subconscious.dev/compare/relace-vs-particle-ai.md)

## Sources

- [Relace](https://relace.ai/)
- [Relace Apply API](https://docs.relace.ai/api-reference/instant-apply/apply)
- [Relace Instant Apply quickstart](https://docs.relace.ai/docs/instant-apply/quickstart)

Pricing and model lineups change often; figures are a snapshot.
