# Baseten vs Luminal

> Baseten serves curated models and custom deployments with the lowest measured TTFT. Luminal compiles models into native kernels for more throughput.

Canonical: https://www.subconscious.dev/compare/baseten-vs-luminal · By The Subconscious Team · Updated September 30, 2026

## How they compare

Baseten's Model APIs cover 13 curated open models with the lowest measured time to first token, and Truss lets teams deploy any model on dedicated GPUs or self-host. Luminal also takes any model, but changes how it runs: a compiler lowers it to primitive ops and emits fused GPU kernels ahead of time, rather than running it through a runtime engine.

Baseten is the mature choice for latency-sensitive production with clear dedicated pricing, about $6.50 an hour for an H100. Luminal targets throughput, reporting 36K tokens per second on GPT-OSS 120B across 8 H100s, and its cloud is early access with rates not published. A team with a custom model could reasonably test both, Baseten for the deployment tooling and Luminal for the engine.

## What each one does

### Baseten

Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.

### Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

## Which is best, and when

### Choose Baseten for

- Lowest time to first token on curated models
- Truss for packaging any model
- Mature dedicated and self-host options

### Choose Luminal for

- Maximum throughput per GPU on a self-chosen model
- Replacing vLLM or TensorRT-LLM with a compiled engine
- An open-source engine teams can run on their own hardware

## At a glance

| Attribute | Baseten | Luminal |
|---|---|---|
| Model access | Open weights, 13 curated | Bring your own weights |
| Flagship models | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | No public catalog |
| Speed | 0.49s TTFT, lowest measured | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | H100 about $6.50/hr dedicated | Pay per use; rates not published |
| Customization | Deploy any model with Truss | Compiles any PyTorch or HF model |
| Deployment | Model APIs, dedicated, self-host | Serverless (early access), on-prem license |
| Long context | Varies by model | - |

## FAQ

### What is the difference between Baseten and Luminal?

Baseten serves curated models and custom deployments with the lowest measured TTFT. Luminal compiles models into native kernels for more throughput.

### When should I choose Baseten over Luminal?

Lowest time to first token on curated models; Truss for packaging any model; Mature dedicated and self-host options.

### When should I choose Luminal over Baseten?

Maximum throughput per GPU on a self-chosen model; Replacing vLLM or TensorRT-LLM with a compiled engine; An open-source engine teams can run on their own hardware.

### Is Baseten or Luminal cheaper?

Baseten: H100 about $6.50/hr dedicated. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Baseten](https://www.subconscious.dev/compare/subconscious-vs-baseten.md), [Subconscious vs Luminal](https://www.subconscious.dev/compare/subconscious-vs-luminal.md).

Full profiles: [Baseten](https://www.subconscious.dev/providers/baseten.md), [Luminal](https://www.subconscious.dev/providers/luminal.md).
