Baseten vs Modal
Both run your own models on GPUs. Modal is general Python compute billed per second; Baseten adds hosted model APIs, record first-token latency and compliance options.
By The Subconscious Team · Updated
Baseten vs Modal: key differences
Modal is serverless compute with GPUs attached. You decorate a Python function, it builds and autoscales the container, and billing runs per second from $0.59 an hour for a T4 to $3.95 for an H100 at list. It has no model catalog and no per-token price. Baseten overlaps on the dedicated side, with Truss packaging any model and billing per GPU minute (an H100 at about $6.50 an hour), but it also runs Model APIs for 13 open models behind OpenAI and Anthropic-compatible endpoints. A team that wants GLM 5.2 or Kimi K3 without writing serving code can call it on Baseten without writing serving code. On Modal that team writes its own vLLM stack.
List price flatters Modal. Non-preemptible US production runs about 3.75x list, near $14.81 an hour for an H100, and keeping containers warm turns a serverless bill into an always-on one. Modal's reach is wider, though: fine-tuning, batch jobs and agent sandboxes all live on one platform, and $30 of free monthly credits helps small teams. Baseten is the tighter fit for production inference with HIPAA, data residency and a 99.99% SLA.
What Baseten and Modal do
Baseten
Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.
Example models: GLM 5.2, gpt-oss 120B
Full Baseten profileModal
Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.
Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper
Full Modal profileShould you choose Baseten or Modal?
Baseten
Choose Baseten for
- Production inference with a 99.99% SLA and HIPAA
- Calling hosted open models without writing serving code
- Model labs wanting a white-label API
Modal
Choose Modal for
- Bursty GPU jobs like embeddings, transcription and batch
- Teams mixing training, inference and sandboxes in Python
- Prototyping on $30 of free monthly credits
Baseten vs Modal at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights, 13 curated | Bring your own weights |
| Flagship models | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | None hosted |
| Speed | 0.49s TTFT, lowest measured | ~1s container boot |
| Price | H100 about $6.50/hr dedicated | Per second; H100 $3.95/hr list |
| Customization | Deploy any model with Truss | Run any training code |
| Deployment | Model APIs, dedicated, self-host | Serverless GPU containers |
| Long context | Varies by model | Depends on the model you deploy |
Frequently asked questions
What is the difference between Baseten and Modal?
Both run your own models on GPUs. Modal is general Python compute billed per second; Baseten adds hosted model APIs, record first-token latency and compliance options.
When should I choose Baseten over Modal?
Production inference with a 99.99% SLA and HIPAA; Calling hosted open models without writing serving code; Model labs wanting a white-label API.
When should I choose Modal over Baseten?
Bursty GPU jobs like embeddings, transcription and batch; Teams mixing training, inference and sandboxes in Python; Prototyping on $30 of free monthly credits.
Is Baseten or Modal cheaper?
Baseten: H100 about $6.50/hr dedicated. Modal: Per second; H100 $3.95/hr list. The cheaper choice depends on the model and workload.
Which has more context, Baseten or Modal?
Baseten: Varies by model. Modal: Depends on the model you deploy.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.