# Modal

> Serverless GPU compute for Python: bring your own model and pay by the second.

Canonical: https://www.subconscious.dev/providers/modal · By The Subconscious Team · Updated September 30, 2026

- Founded: 2021
- Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper
- Website: https://modal.com

## Overview

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Modal sells compute, so it has no model catalog and no per-token price. Teams bring their own weights and serving code, which makes it a strong fit for custom models, fine-tunes, embeddings, OCR and batch jobs. Container boot takes about one second; the real cold start is loading weights, which baking them into the image and memory snapshotting both shrink. The free Starter plan renews $30 of credits every month, the most generous no-commitment allowance in the category.

## Upsides

- Python-native developer experience that ships GPU services in an afternoon.
- Per-second billing beats reserved GPUs on bursty load below roughly 80% utilization.
- Covers inference, fine-tuning, batch jobs and agent sandboxes on one platform.

## Core use cases

- Spiky or queued GPU work like embeddings, reranking, transcription and media jobs.
- Dedicated serverless endpoints for private or fine-tuned models.

## Downsides

- Non-preemptible US production runs about 3.75x list price, putting an H100 near $14.81 an hour.
- Keeping containers warm to avoid cold starts turns the serverless bill into an always-on one.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Bring your own weights |
| Flagship models | None hosted |
| Speed | ~1s container boot |
| Price | Per second; H100 $3.95/hr list |
| Customization | Run any training code |
| Deployment | Serverless GPU containers |
| Long context | Depends on the model you deploy |

## FAQ

### What is Modal?

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

### What is Modal best for?

Spiky or queued GPU work like embeddings, reranking, transcription and media jobs; Dedicated serverless endpoints for private or fine-tuned models.

### How much does Modal cost?

Modal pricing at a glance: Per second; H100 $3.95/hr list. Rates change often, so check Modal's pricing page before committing.

### How much context does Modal support?

Modal's long-context support: Depends on the model you deploy.

### What are the downsides of Modal?

Non-preemptible US production runs about 3.75x list price, putting an H100 near $14.81 an hour; Keeping containers warm to avoid cold starts turns the serverless bill into an always-on one.

### What are the best alternatives to Modal?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Modal on this site.

## Comparisons

- [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md)
- [OpenAI vs Modal](https://www.subconscious.dev/compare/openai-vs-modal.md)
- [Anthropic vs Modal](https://www.subconscious.dev/compare/anthropic-vs-modal.md)
- [Google Vertex AI vs Modal](https://www.subconscious.dev/compare/google-vertex-vs-modal.md)
- [Amazon Bedrock vs Modal](https://www.subconscious.dev/compare/aws-bedrock-vs-modal.md)
- [Together AI vs Modal](https://www.subconscious.dev/compare/together-ai-vs-modal.md)
- [Fireworks AI vs Modal](https://www.subconscious.dev/compare/fireworks-vs-modal.md)
- [Baseten vs Modal](https://www.subconscious.dev/compare/baseten-vs-modal.md)
- [Groq vs Modal](https://www.subconscious.dev/compare/groq-vs-modal.md)
- [Cerebras vs Modal](https://www.subconscious.dev/compare/cerebras-vs-modal.md)
- [DeepInfra vs Modal](https://www.subconscious.dev/compare/deepinfra-vs-modal.md)
- [Modal vs xAI](https://www.subconscious.dev/compare/modal-vs-xai.md)
- [Modal vs DeepSeek](https://www.subconscious.dev/compare/modal-vs-deepseek.md)
- [Modal vs Moonshot AI](https://www.subconscious.dev/compare/modal-vs-moonshot-ai.md)
- [Modal vs Z.ai](https://www.subconscious.dev/compare/modal-vs-z-ai.md)
- [Modal vs Alibaba Cloud](https://www.subconscious.dev/compare/modal-vs-alibaba-cloud.md)
- [Modal vs Meta](https://www.subconscious.dev/compare/modal-vs-meta.md)
- [Modal vs SambaNova](https://www.subconscious.dev/compare/modal-vs-sambanova.md)
- [Modal vs Nebius](https://www.subconscious.dev/compare/modal-vs-nebius.md)
- [Modal vs fal](https://www.subconscious.dev/compare/modal-vs-fal.md)
- [Modal vs Novita AI](https://www.subconscious.dev/compare/modal-vs-novita-ai.md)
- [Modal vs Parasail](https://www.subconscious.dev/compare/modal-vs-parasail.md)
- [Modal vs Inference.net](https://www.subconscious.dev/compare/modal-vs-inference-net.md)
- [Modal vs GMI Cloud](https://www.subconscious.dev/compare/modal-vs-gmi-cloud.md)
- [Modal vs Sail Research](https://www.subconscious.dev/compare/modal-vs-sail-research.md)
- [Modal vs Morph](https://www.subconscious.dev/compare/modal-vs-morph.md)
- [Modal vs Relace](https://www.subconscious.dev/compare/modal-vs-relace.md)
- [Modal vs TypeSafe AI](https://www.subconscious.dev/compare/modal-vs-typesafe-ai.md)
- [Modal vs StepFun](https://www.subconscious.dev/compare/modal-vs-stepfun.md)
- [Modal vs Runware](https://www.subconscious.dev/compare/modal-vs-runware.md)
- [Modal vs StreamLake](https://www.subconscious.dev/compare/modal-vs-streamlake.md)
- [Modal vs Wafer](https://www.subconscious.dev/compare/modal-vs-wafer.md)
- [Modal vs RunInfra](https://www.subconscious.dev/compare/modal-vs-runinfra.md)
- [Modal vs Particle.AI](https://www.subconscious.dev/compare/modal-vs-particle-ai.md)

## Sources

- [Modal review, Continuum](https://continuumcode.ai/guides/modal-review/)
- [Modal review, Alatirok](https://alatirok.com/modal-labs-review-per-second-gpu-billing/)
- [Modal review, HostFleet](https://hostfleet.net/modal-for-ai-inference-apis-and-jobs/)

Pricing and model lineups change often; figures are a snapshot.
