Modal vs Luminal
Modal runs your Python on serverless GPUs by the second. Luminal compiles your model so each GPU serves more tokens.
By The Subconscious Team · Updated
Modal vs Luminal: key differences
Modal is general serverless GPU compute: write Python, get containers that boot in about a second, and pay per second, with an H100 at $3.95 an hour list. You pick the inference engine yourself. Luminal is that engine layer. Its compiler turns a model into native kernels ahead of time, and the company also sells serverless endpoints with scale to zero and an on-prem license.
The two are complements more than rivals. Modal gives flexibility for training, batch jobs and any code; Luminal gives throughput, reporting 36K tokens per second on GPT-OSS 120B over 8 H100s against 26K for vLLM. Teams on Modal with large inference bills could test Luminal's open-source compiler inside their own containers.
What Modal and Luminal do
Modal
Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.
Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper
Full Modal profileLuminal
Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.
Example models: GPT-OSS 120B, Llama 3 8B
Full Luminal profileShould you choose Modal or Luminal?
Modal
Choose Modal for
- Arbitrary Python and training jobs on GPUs
- Per-second billing
- Full control of the serving stack
Luminal
Choose Luminal for
- Maximum throughput per GPU on a self-chosen model
- Replacing vLLM or TensorRT-LLM with a compiled engine
- Managed endpoints that scale to zero without writing infra code
Modal vs Luminal at a glance
| Attribute | ||
|---|---|---|
| Model access | Bring your own weights | Bring your own weights |
| Flagship models | None hosted | No public catalog |
| Speed | ~1s container boot | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | Per second; H100 $3.95/hr list | Pay per use; rates not published |
| Customization | Run any training code | Compiles any PyTorch or HF model |
| Deployment | Serverless GPU containers | Serverless (early access), on-prem license |
| Long context | Depends on the model you deploy | Unknown |
Frequently asked questions
What is the difference between Modal and Luminal?
Modal runs your Python on serverless GPUs by the second. Luminal compiles your model so each GPU serves more tokens.
When should I choose Modal over Luminal?
Arbitrary Python and training jobs on GPUs; Per-second billing; Full control of the serving stack.
When should I choose Luminal over Modal?
Maximum throughput per GPU on a self-chosen model; Replacing vLLM or TensorRT-LLM with a compiled engine; Managed endpoints that scale to zero without writing infra code.
Is Modal or Luminal cheaper?
Modal: Per second; H100 $3.95/hr list. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.