# Modal vs Relace

> Relace provides fast specialist models for coding agents: apply, search and compaction. Modal provides the GPUs and sandboxes an agent runs on. Different layers of the same stack.

Canonical: https://www.subconscious.dev/compare/modal-vs-relace · By The Subconscious Team · Updated September 30, 2026

## How they compare

Relace trains small models that act as tools for coding agents. relace-apply-3 merges edit snippets at about 10,000 tokens per second with 128K tokens of input and output, its agentic search explores large codebases in parallel, and a compaction model runs at 50,000 tokens per second. It offers a hosted API or self-hosted deployment with guided onboarding. Modal is general serverless compute: GPUs from T4 to B300, per-second billing, and sandboxes for agents. Relace sells finished tools; Modal sells the machine time to run your own.

Teams would combine them rather than choose. Modal runs the agent's containers and any custom models, and Relace handles merges and retrieval, taking work off expensive frontier models. Relace errors past 128K tokens, so very large files need a fallback model, which could itself run on Modal. For enterprises keeping code in-house, Relace's self-hosted option matters. Modal's costs to watch are cold starts from loading weights and about 3.75x list for non-preemptible US production.

## What each one does

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

### Relace

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

## Which is best, and when

### Choose Modal for

- The compute layer for coding-agent sandboxes.
- Serving fallback or custom models on demand.
- Per-second billing for spiky CI workloads.

### Choose Relace for

- Instant apply at about 10,000 tokens per second.
- Parallel agentic search over large repos.
- Self-hosted coding utilities for private code.

## At a glance

| Attribute | Modal | Relace |
|---|---|---|
| Model access | Bring your own weights | Specialist models |
| Flagship models | None hosted | relace-apply-3, agentic search |
| Speed | ~1s container boot | ~10,000 tok/s apply |
| Price | Per second; H100 $3.95/hr list | 3x+ cheaper than full rewrites |
| Customization | Run any training code | - |
| Deployment | Serverless GPU containers | Hosted API or self-hosted |
| Long context | Depends on the model you deploy | 128K max |

## FAQ

### What is the difference between Modal and Relace?

Relace provides fast specialist models for coding agents: apply, search and compaction. Modal provides the GPUs and sandboxes an agent runs on. Different layers of the same stack.

### When should I choose Modal over Relace?

The compute layer for coding-agent sandboxes; Serving fallback or custom models on demand; Per-second billing for spiky CI workloads.

### When should I choose Relace over Modal?

Instant apply at about 10,000 tokens per second; Parallel agentic search over large repos; Self-hosted coding utilities for private code.

### Is Modal or Relace cheaper?

Modal: Per second; H100 $3.95/hr list. Relace: 3x+ cheaper than full rewrites. The cheaper choice depends on the model and workload.

### Which has more context, Modal or Relace?

Modal: Depends on the model you deploy. Relace: 128K max.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md), [Subconscious vs Relace](https://www.subconscious.dev/compare/subconscious-vs-relace.md).

Full profiles: [Modal](https://www.subconscious.dev/providers/modal.md), [Relace](https://www.subconscious.dev/providers/relace.md).
