# Modal vs Luminal

> Modal runs your Python on serverless GPUs by the second. Luminal compiles your model so each GPU serves more tokens.

Canonical: https://www.subconscious.dev/compare/modal-vs-luminal · By The Subconscious Team · Updated September 30, 2026

## How they compare

Modal is general serverless GPU compute: write Python, get containers that boot in about a second, and pay per second, with an H100 at $3.95 an hour list. You pick the inference engine yourself. Luminal is that engine layer. Its compiler turns a model into native kernels ahead of time, and the company also sells serverless endpoints with scale to zero and an on-prem license.

The two are complements more than rivals. Modal gives flexibility for training, batch jobs and any code; Luminal gives throughput, reporting 36K tokens per second on GPT-OSS 120B over 8 H100s against 26K for vLLM. Teams on Modal with large inference bills could test Luminal's open-source compiler inside their own containers.

## What each one does

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

### Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

## Which is best, and when

### Choose Modal for

- Arbitrary Python and training jobs on GPUs
- Per-second billing
- Full control of the serving stack

### Choose Luminal for

- Maximum throughput per GPU on a self-chosen model
- Replacing vLLM or TensorRT-LLM with a compiled engine
- Managed endpoints that scale to zero without writing infra code

## At a glance

| Attribute | Modal | Luminal |
|---|---|---|
| Model access | Bring your own weights | Bring your own weights |
| Flagship models | None hosted | No public catalog |
| Speed | ~1s container boot | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | Per second; H100 $3.95/hr list | Pay per use; rates not published |
| Customization | Run any training code | Compiles any PyTorch or HF model |
| Deployment | Serverless GPU containers | Serverless (early access), on-prem license |
| Long context | Depends on the model you deploy | - |

## FAQ

### What is the difference between Modal and Luminal?

Modal runs your Python on serverless GPUs by the second. Luminal compiles your model so each GPU serves more tokens.

### When should I choose Modal over Luminal?

Arbitrary Python and training jobs on GPUs; Per-second billing; Full control of the serving stack.

### When should I choose Luminal over Modal?

Maximum throughput per GPU on a self-chosen model; Replacing vLLM or TensorRT-LLM with a compiled engine; Managed endpoints that scale to zero without writing infra code.

### Is Modal or Luminal cheaper?

Modal: Per second; H100 $3.95/hr list. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md), [Subconscious vs Luminal](https://www.subconscious.dev/compare/subconscious-vs-luminal.md).

Full profiles: [Modal](https://www.subconscious.dev/providers/modal.md), [Luminal](https://www.subconscious.dev/providers/luminal.md).
