# Baseten vs Modal

> Both run your own models on GPUs. Modal is general Python compute billed per second; Baseten adds hosted model APIs, record first-token latency and compliance options.

Canonical: https://www.subconscious.dev/compare/baseten-vs-modal · By The Subconscious Team · Updated September 30, 2026

## How they compare

Modal is serverless compute with GPUs attached. You decorate a Python function, it builds and autoscales the container, and billing runs per second from $0.59 an hour for a T4 to $3.95 for an H100 at list. It has no model catalog and no per-token price. Baseten overlaps on the dedicated side, with Truss packaging any model and billing per GPU minute (an H100 at about $6.50 an hour), but it also runs Model APIs for 13 open models behind OpenAI and Anthropic-compatible endpoints. A team that wants GLM 5.2 or Kimi K3 without writing serving code can call it on Baseten without writing serving code. On Modal that team writes its own vLLM stack.

List price flatters Modal. Non-preemptible US production runs about 3.75x list, near $14.81 an hour for an H100, and keeping containers warm turns a serverless bill into an always-on one. Modal's reach is wider, though: fine-tuning, batch jobs and agent sandboxes all live on one platform, and $30 of free monthly credits helps small teams. Baseten is the tighter fit for production inference with HIPAA, data residency and a 99.99% SLA.

## What each one does

### Baseten

Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

## Which is best, and when

### Choose Baseten for

- Production inference with a 99.99% SLA and HIPAA
- Calling hosted open models without writing serving code
- Model labs wanting a white-label API

### Choose Modal for

- Bursty GPU jobs like embeddings, transcription and batch
- Teams mixing training, inference and sandboxes in Python
- Prototyping on $30 of free monthly credits

## At a glance

| Attribute | Baseten | Modal |
|---|---|---|
| Model access | Open weights, 13 curated | Bring your own weights |
| Flagship models | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | None hosted |
| Speed | 0.49s TTFT, lowest measured | ~1s container boot |
| Price | H100 about $6.50/hr dedicated | Per second; H100 $3.95/hr list |
| Customization | Deploy any model with Truss | Run any training code |
| Deployment | Model APIs, dedicated, self-host | Serverless GPU containers |
| Long context | Varies by model | Depends on the model you deploy |

## FAQ

### What is the difference between Baseten and Modal?

Both run your own models on GPUs. Modal is general Python compute billed per second; Baseten adds hosted model APIs, record first-token latency and compliance options.

### When should I choose Baseten over Modal?

Production inference with a 99.99% SLA and HIPAA; Calling hosted open models without writing serving code; Model labs wanting a white-label API.

### When should I choose Modal over Baseten?

Bursty GPU jobs like embeddings, transcription and batch; Teams mixing training, inference and sandboxes in Python; Prototyping on $30 of free monthly credits.

### Is Baseten or Modal cheaper?

Baseten: H100 about $6.50/hr dedicated. Modal: Per second; H100 $3.95/hr list. The cheaper choice depends on the model and workload.

### Which has more context, Baseten or Modal?

Baseten: Varies by model. Modal: Depends on the model you deploy.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Baseten](https://www.subconscious.dev/compare/subconscious-vs-baseten.md), [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md).

Full profiles: [Baseten](https://www.subconscious.dev/providers/baseten.md), [Modal](https://www.subconscious.dev/providers/modal.md).
