vs

Baseten vs Modal

Both run your own models on GPUs. Modal is general Python compute billed per second; Baseten adds hosted model APIs, record first-token latency and compliance options.

By The Subconscious Team · Updated

Baseten vs Modal: key differences

Modal is serverless compute with GPUs attached. You decorate a Python function, it builds and autoscales the container, and billing runs per second from $0.59 an hour for a T4 to $3.95 for an H100 at list. It has no model catalog and no per-token price. Baseten overlaps on the dedicated side, with Truss packaging any model and billing per GPU minute (an H100 at about $6.50 an hour), but it also runs Model APIs for 13 open models behind OpenAI and Anthropic-compatible endpoints. A team that wants GLM 5.2 or Kimi K3 without writing serving code can call it on Baseten without writing serving code. On Modal that team writes its own vLLM stack.

List price flatters Modal. Non-preemptible US production runs about 3.75x list, near $14.81 an hour for an H100, and keeping containers warm turns a serverless bill into an always-on one. Modal's reach is wider, though: fine-tuning, batch jobs and agent sandboxes all live on one platform, and $30 of free monthly credits helps small teams. Baseten is the tighter fit for production inference with HIPAA, data residency and a 99.99% SLA.

What Baseten and Modal do

Baseten

Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.

Example models: GLM 5.2, gpt-oss 120B

Full Baseten profile

Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper

Full Modal profile

Should you choose Baseten or Modal?

Baseten

Choose Baseten for

  • Production inference with a 99.99% SLA and HIPAA
  • Calling hosted open models without writing serving code
  • Model labs wanting a white-label API

Modal

Choose Modal for

  • Bursty GPU jobs like embeddings, transcription and batch
  • Teams mixing training, inference and sandboxes in Python
  • Prototyping on $30 of free monthly credits

Baseten vs Modal at a glance

AttributeBasetenModal
Model accessOpen weights, 13 curatedBring your own weights
Flagship modelsGLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120BNone hosted
Speed0.49s TTFT, lowest measured~1s container boot
PriceH100 about $6.50/hr dedicatedPer second; H100 $3.95/hr list
CustomizationDeploy any model with TrussRun any training code
DeploymentModel APIs, dedicated, self-hostServerless GPU containers
Long contextVaries by modelDepends on the model you deploy

Frequently asked questions

What is the difference between Baseten and Modal?

Both run your own models on GPUs. Modal is general Python compute billed per second; Baseten adds hosted model APIs, record first-token latency and compliance options.

When should I choose Baseten over Modal?

Production inference with a 99.99% SLA and HIPAA; Calling hosted open models without writing serving code; Model labs wanting a white-label API.

When should I choose Modal over Baseten?

Bursty GPU jobs like embeddings, transcription and batch; Teams mixing training, inference and sandboxes in Python; Prototyping on $30 of free monthly credits.

Is Baseten or Modal cheaper?

Baseten: H100 about $6.50/hr dedicated. Modal: Per second; H100 $3.95/hr list. The cheaper choice depends on the model and workload.

Which has more context, Baseten or Modal?

Baseten: Varies by model. Modal: Depends on the model you deploy.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.