# Sail Research vs Relace

> Sail Research runs the reasoning model cheaply by making it wait. Relace runs the utility steps (apply, search, compaction) as fast as it can. They split one coding agent's work.

Canonical: https://www.subconscious.dev/compare/sail-research-vs-relace · By The Subconscious Team · Updated September 30, 2026

## How they compare

Relace trains small models that act as tools inside coding agents. relace-apply-3 merges lazy edit snippets at about 10,000 tokens per second, its agentic search explores large codebases in seconds, and a compaction model runs at 50,000 tokens per second. Sail Research hosts the general open models an agent thinks with, such as Kimi K2.6, GLM-5, GPT-OSS 120B and Qwen 3.6, and prices them by patience. The priority window targets about a one-minute turn at roughly 30 to 50% off, and the standard window about five minutes at 45 to 65% off.

Neither replaces the other. Relace has no general-purpose model serving, and Sail has no specialist apply or search models. A PR review or automated-fix pipeline could run its planner on Sail's standard window and call Relace for retrieval and merges, keeping expensive tokens off the main model. Limits differ: Relace returns an error past 128K tokens, so very large files need a fallback, while Sail is explicitly unsuited to anything interactive. Relace also offers self-hosted deployment for teams that keep code in-house.

## What each one does

### Sail Research

Sail Research sells throughput over latency. Founders Neil Movva and Samir Menon built a serving stack that packs as much work as possible into every GPU, and customers state how long they can wait through completion windows. The priority window targets about a one-minute turn for roughly 30 to 50% off the immediate asap price. The default standard window targets about five minutes for 45 to 65% off. The flex window runs off-peak for 60 to 80% off.

### Relace

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

## Which is best, and when

### Choose Sail Research for

- The main reasoning model for unattended coding agents.
- Cheap open-model tokens when turns can take minutes.
- Customer LoRA fine-tunes served over OpenAI or Anthropic APIs.

### Choose Relace for

- Fast file merges and repo search inside an agent loop.
- Context compaction at 50,000 tokens per second.
- Self-hosted coding utilities for code that stays in-house.

## At a glance

| Attribute | Sail Research | Relace |
|---|---|---|
| Model access | Open weights | Specialist models |
| Flagship models | Kimi K2.6, GLM-5, GPT-OSS 120B | relace-apply-3, agentic search |
| Speed | Minutes per turn by design | ~10,000 tok/s apply |
| Price | 30–80% off by completion window | 3x+ cheaper than full rewrites |
| Customization | Customer LoRA fine-tunes | - |
| Deployment | API plus Sailboxes | Hosted API or self-hosted |
| Long context | Varies by model | 128K max |

## FAQ

### What is the difference between Sail Research and Relace?

Sail Research runs the reasoning model cheaply by making it wait. Relace runs the utility steps (apply, search, compaction) as fast as it can. They split one coding agent's work.

### When should I choose Sail Research over Relace?

The main reasoning model for unattended coding agents; Cheap open-model tokens when turns can take minutes; Customer LoRA fine-tunes served over OpenAI or Anthropic APIs.

### When should I choose Relace over Sail Research?

Fast file merges and repo search inside an agent loop; Context compaction at 50,000 tokens per second; Self-hosted coding utilities for code that stays in-house.

### Is Sail Research or Relace cheaper?

Sail Research: 30–80% off by completion window. Relace: 3x+ cheaper than full rewrites. The cheaper choice depends on the model and workload.

### Which has more context, Sail Research or Relace?

Sail Research: Varies by model. Relace: 128K max.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Sail Research](https://www.subconscious.dev/compare/subconscious-vs-sail-research.md), [Subconscious vs Relace](https://www.subconscious.dev/compare/subconscious-vs-relace.md).

Full profiles: [Sail Research](https://www.subconscious.dev/providers/sail-research.md), [Relace](https://www.subconscious.dev/providers/relace.md).
