# Modal vs Sail Research

> Sail Research sells discounted open-model tokens for patient, long-running agents. Modal sells per-second GPUs for your own code. Both offer agent sandboxes, with very different pricing.

Canonical: https://www.subconscious.dev/compare/modal-vs-sail-research · By The Subconscious Team · Updated September 30, 2026

## How they compare

Sail Research is an inference host tuned for throughput. It serves open models like Kimi K2.6, GLM-5 and GPT-OSS 120B plus customer LoRAs, and discounts 30 to 80% off its asap price depending on the completion window, from a one-minute priority turn to off-peak flex. Its Sailboxes give agents persistent compute that can run indefinitely. Modal is general serverless GPU compute with per-second billing and no catalog, and it also covers agent sandboxes, fine-tuning and batch jobs. Sail sells finished tokens, Modal sells the hardware time underneath.

For background agents on popular open models, Sail usually costs less and needs no serving code, since the tokens and the sandbox come together. Modal fits when the model is custom, when the job is not text generation, or when the work is interactive, which Sail explicitly does not serve. Modal's bursty billing pays off below roughly 80% utilization, but warm containers turn it always-on. Sail's 3x to 10x savings claim is its own, and it offers open models only.

## What each one does

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

### Sail Research

Sail Research sells throughput over latency. Founders Neil Movva and Samir Menon built a serving stack that packs as much work as possible into every GPU, and customers state how long they can wait through completion windows. The priority window targets about a one-minute turn for roughly 30 to 50% off the immediate asap price. The default standard window targets about five minutes for 45 to 65% off. The flex window runs off-peak for 60 to 80% off.

## Which is best, and when

### Choose Modal for

- Custom models and non-LLM GPU work.
- Interactive endpoints that scale to zero.
- Python-native control of training and serving.

### Choose Sail Research for

- Hours-long background agents on open models.
- Deep discounts for workloads that can wait minutes.
- Hosted tokens plus persistent agent sandboxes together.

## At a glance

| Attribute | Modal | Sail Research |
|---|---|---|
| Model access | Bring your own weights | Open weights |
| Flagship models | None hosted | Kimi K2.6, GLM-5, GPT-OSS 120B |
| Speed | ~1s container boot | Minutes per turn by design |
| Price | Per second; H100 $3.95/hr list | 30–80% off by completion window |
| Customization | Run any training code | Customer LoRA fine-tunes |
| Deployment | Serverless GPU containers | API plus Sailboxes |
| Long context | Depends on the model you deploy | Varies by model |

## FAQ

### What is the difference between Modal and Sail Research?

Sail Research sells discounted open-model tokens for patient, long-running agents. Modal sells per-second GPUs for your own code. Both offer agent sandboxes, with very different pricing.

### When should I choose Modal over Sail Research?

Custom models and non-LLM GPU work; Interactive endpoints that scale to zero; Python-native control of training and serving.

### When should I choose Sail Research over Modal?

Hours-long background agents on open models; Deep discounts for workloads that can wait minutes; Hosted tokens plus persistent agent sandboxes together.

### Is Modal or Sail Research cheaper?

Modal: Per second; H100 $3.95/hr list. Sail Research: 30–80% off by completion window. The cheaper choice depends on the model and workload.

### Which has more context, Modal or Sail Research?

Modal: Depends on the model you deploy. Sail Research: Varies by model.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md), [Subconscious vs Sail Research](https://www.subconscious.dev/compare/subconscious-vs-sail-research.md).

Full profiles: [Modal](https://www.subconscious.dev/providers/modal.md), [Sail Research](https://www.subconscious.dev/providers/sail-research.md).
