# Luminal

> An open-source compiler that turns models into native GPU code, sold as serverless or on-prem.

Canonical: https://www.subconscious.dev/providers/luminal · By The Subconscious Team · Updated September 30, 2026

- Founded: 2025
- Example models: GPT-OSS 120B, Llama 3 8B
- Website: https://www.luminal.com

## Overview

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

The company sells that compiler two ways. Luminal Cloud is serverless inference in early access: upload a PyTorch or Hugging Face model, get a compiled endpoint with automatic batching and scale to zero, and pay per use. The on-prem license adds custom kernel work, dedicated engineers and SLAs. Luminal reports GPT-OSS 120B at 36K tokens per second on 8 H100s, against 28K for TensorRT-LLM and 26K for vLLM. Founders Joe Fioti, Jake Stevens and Matthew Gunton came from Intel, Apple and Amazon, went through Y Combinator in Summer 2025, and raised a $5.3M seed led by Felicis in November 2025.

## Upsides

- Ahead-of-time compilation can beat runtime engines on throughput for the same model and hardware.
- Open-source compiler, so teams can inspect it or run it themselves before buying.
- Brings your own model, including custom architectures that hosted catalogs skip.

## Core use cases

- High-throughput serving of a custom or fine-tuned model on owned GPUs.
- Teams that want an inference engine faster than vLLM without hand-writing kernels.

## Downsides

- Very young company; the cloud is early access with no public price list or model catalog.
- Speed figures are self-reported aggregate throughput, not independent per-request benchmarks.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Bring your own weights |
| Flagship models | No public catalog |
| Speed | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | Pay per use; rates not published |
| Customization | Compiles any PyTorch or HF model |
| Deployment | Serverless (early access), on-prem license |
| Long context | - |

## FAQ

### What is Luminal?

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

### What is Luminal best for?

High-throughput serving of a custom or fine-tuned model on owned GPUs; Teams that want an inference engine faster than vLLM without hand-writing kernels.

### How much does Luminal cost?

Luminal pricing at a glance: Pay per use; rates not published. Rates change often, so check Luminal's pricing page before committing.

### What are the downsides of Luminal?

Very young company; the cloud is early access with no public price list or model catalog; Speed figures are self-reported aggregate throughput, not independent per-request benchmarks.

### What are the best alternatives to Luminal?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Luminal on this site.

## Comparisons

- [Subconscious vs Luminal](https://www.subconscious.dev/compare/subconscious-vs-luminal.md)
- [OpenAI vs Luminal](https://www.subconscious.dev/compare/openai-vs-luminal.md)
- [Anthropic vs Luminal](https://www.subconscious.dev/compare/anthropic-vs-luminal.md)
- [Google Vertex AI vs Luminal](https://www.subconscious.dev/compare/google-vertex-vs-luminal.md)
- [Amazon Bedrock vs Luminal](https://www.subconscious.dev/compare/aws-bedrock-vs-luminal.md)
- [Together AI vs Luminal](https://www.subconscious.dev/compare/together-ai-vs-luminal.md)
- [Fireworks AI vs Luminal](https://www.subconscious.dev/compare/fireworks-vs-luminal.md)
- [Baseten vs Luminal](https://www.subconscious.dev/compare/baseten-vs-luminal.md)
- [Groq vs Luminal](https://www.subconscious.dev/compare/groq-vs-luminal.md)
- [Cerebras vs Luminal](https://www.subconscious.dev/compare/cerebras-vs-luminal.md)
- [DeepInfra vs Luminal](https://www.subconscious.dev/compare/deepinfra-vs-luminal.md)
- [Hugging Face Inference Providers vs Luminal](https://www.subconscious.dev/compare/hugging-face-vs-luminal.md)
- [Modal vs Luminal](https://www.subconscious.dev/compare/modal-vs-luminal.md)
- [Cloudflare Workers AI vs Luminal](https://www.subconscious.dev/compare/cloudflare-workers-ai-vs-luminal.md)
- [xAI vs Luminal](https://www.subconscious.dev/compare/xai-vs-luminal.md)
- [Mistral AI vs Luminal](https://www.subconscious.dev/compare/mistral-ai-vs-luminal.md)
- [DeepSeek vs Luminal](https://www.subconscious.dev/compare/deepseek-vs-luminal.md)
- [Moonshot AI vs Luminal](https://www.subconscious.dev/compare/moonshot-ai-vs-luminal.md)
- [Z.ai vs Luminal](https://www.subconscious.dev/compare/z-ai-vs-luminal.md)
- [Alibaba Cloud vs Luminal](https://www.subconscious.dev/compare/alibaba-cloud-vs-luminal.md)
- [Meta vs Luminal](https://www.subconscious.dev/compare/meta-vs-luminal.md)
- [Cohere vs Luminal](https://www.subconscious.dev/compare/cohere-vs-luminal.md)
- [SambaNova vs Luminal](https://www.subconscious.dev/compare/sambanova-vs-luminal.md)
- [Nebius vs Luminal](https://www.subconscious.dev/compare/nebius-vs-luminal.md)
- [Crusoe vs Luminal](https://www.subconscious.dev/compare/crusoe-vs-luminal.md)
- [fal vs Luminal](https://www.subconscious.dev/compare/fal-vs-luminal.md)
- [Novita AI vs Luminal](https://www.subconscious.dev/compare/novita-ai-vs-luminal.md)
- [Venice vs Luminal](https://www.subconscious.dev/compare/venice-vs-luminal.md)
- [Parasail vs Luminal](https://www.subconscious.dev/compare/parasail-vs-luminal.md)
- [Inference.net vs Luminal](https://www.subconscious.dev/compare/inference-net-vs-luminal.md)
- [GMI Cloud vs Luminal](https://www.subconscious.dev/compare/gmi-cloud-vs-luminal.md)
- [Thinking Machines vs Luminal](https://www.subconscious.dev/compare/thinking-machines-vs-luminal.md)
- [Sail Research vs Luminal](https://www.subconscious.dev/compare/sail-research-vs-luminal.md)
- [Morph vs Luminal](https://www.subconscious.dev/compare/morph-vs-luminal.md)
- [Relace vs Luminal](https://www.subconscious.dev/compare/relace-vs-luminal.md)
- [TypeSafe AI vs Luminal](https://www.subconscious.dev/compare/typesafe-ai-vs-luminal.md)
- [StepFun vs Luminal](https://www.subconscious.dev/compare/stepfun-vs-luminal.md)
- [Runware vs Luminal](https://www.subconscious.dev/compare/runware-vs-luminal.md)
- [StreamLake vs Luminal](https://www.subconscious.dev/compare/streamlake-vs-luminal.md)
- [Wafer vs Luminal](https://www.subconscious.dev/compare/wafer-vs-luminal.md)
- [RunInfra vs Luminal](https://www.subconscious.dev/compare/runinfra-vs-luminal.md)
- [Particle.AI vs Luminal](https://www.subconscious.dev/compare/particle-ai-vs-luminal.md)
- [Luminal vs Infron](https://www.subconscious.dev/compare/luminal-vs-infron.md)

## Sources

- [Luminal](https://www.luminal.com/)
- [Luminal on GitHub](https://github.com/luminal-ai/luminal)
- [Luminal seed round, TechCrunch](https://techcrunch.com/2025/11/17/luminal-raises-5-3-million-to-build-a-better-gpu-code-framework/)

Pricing and model lineups change often; figures are a snapshot.
