Long-running agents deserve better inference.
vs

Modal vs Luminal

Modal runs your Python on serverless GPUs by the second. Luminal compiles your model so each GPU serves more tokens.

By The Subconscious Team · Updated

Modal vs Luminal: key differences

Modal is general serverless GPU compute: write Python, get containers that boot in about a second, and pay per second, with an H100 at $3.95 an hour list. You pick the inference engine yourself. Luminal is that engine layer. Its compiler turns a model into native kernels ahead of time, and the company also sells serverless endpoints with scale to zero and an on-prem license.

The two are complements more than rivals. Modal gives flexibility for training, batch jobs and any code; Luminal gives throughput, reporting 36K tokens per second on GPT-OSS 120B over 8 H100s against 26K for vLLM. Teams on Modal with large inference bills could test Luminal's open-source compiler inside their own containers.

What Modal and Luminal do

Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper

Full Modal profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Modal or Luminal?

Modal

Choose Modal for

  • Arbitrary Python and training jobs on GPUs
  • Per-second billing
  • Full control of the serving stack

Luminal

Choose Luminal for

  • Maximum throughput per GPU on a self-chosen model
  • Replacing vLLM or TensorRT-LLM with a compiled engine
  • Managed endpoints that scale to zero without writing infra code

Modal vs Luminal at a glance

AttributeModalLuminal
Model accessBring your own weightsBring your own weights
Flagship modelsNone hostedNo public catalog
Speed~1s container boot36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
PricePer second; H100 $3.95/hr listPay per use; rates not published
CustomizationRun any training codeCompiles any PyTorch or HF model
DeploymentServerless GPU containersServerless (early access), on-prem license
Long contextDepends on the model you deployUnknown

Frequently asked questions

What is the difference between Modal and Luminal?

Modal runs your Python on serverless GPUs by the second. Luminal compiles your model so each GPU serves more tokens.

When should I choose Modal over Luminal?

Arbitrary Python and training jobs on GPUs; Per-second billing; Full control of the serving stack.

When should I choose Luminal over Modal?

Maximum throughput per GPU on a self-chosen model; Replacing vLLM or TensorRT-LLM with a compiled engine; Managed endpoints that scale to zero without writing infra code.

Is Modal or Luminal cheaper?

Modal: Per second; H100 $3.95/hr list. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.