Long-running agents deserve better inference.
vs

Baseten vs Luminal

Baseten serves curated models and custom deployments with the lowest measured TTFT. Luminal compiles models into native kernels for more throughput.

By The Subconscious Team · Updated

Baseten vs Luminal: key differences

Baseten's Model APIs cover 13 curated open models with the lowest measured time to first token, and Truss lets teams deploy any model on dedicated GPUs or self-host. Luminal also takes any model, but changes how it runs: a compiler lowers it to primitive ops and emits fused GPU kernels ahead of time, rather than running it through a runtime engine.

Baseten is the mature choice for latency-sensitive production with clear dedicated pricing, about $6.50 an hour for an H100. Luminal targets throughput, reporting 36K tokens per second on GPT-OSS 120B across 8 H100s, and its cloud is early access with rates not published. A team with a custom model could reasonably test both, Baseten for the deployment tooling and Luminal for the engine.

What Baseten and Luminal do

Baseten

Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.

Example models: GLM 5.2, gpt-oss 120B

Full Baseten profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Baseten or Luminal?

Baseten

Choose Baseten for

  • Lowest time to first token on curated models
  • Truss for packaging any model
  • Mature dedicated and self-host options

Luminal

Choose Luminal for

  • Maximum throughput per GPU on a self-chosen model
  • Replacing vLLM or TensorRT-LLM with a compiled engine
  • An open-source engine teams can run on their own hardware

Baseten vs Luminal at a glance

AttributeBasetenLuminal
Model accessOpen weights, 13 curatedBring your own weights
Flagship modelsGLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120BNo public catalog
Speed0.49s TTFT, lowest measured36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
PriceH100 about $6.50/hr dedicatedPay per use; rates not published
CustomizationDeploy any model with TrussCompiles any PyTorch or HF model
DeploymentModel APIs, dedicated, self-hostServerless (early access), on-prem license
Long contextVaries by modelUnknown

Frequently asked questions

What is the difference between Baseten and Luminal?

Baseten serves curated models and custom deployments with the lowest measured TTFT. Luminal compiles models into native kernels for more throughput.

When should I choose Baseten over Luminal?

Lowest time to first token on curated models; Truss for packaging any model; Mature dedicated and self-host options.

When should I choose Luminal over Baseten?

Maximum throughput per GPU on a self-chosen model; Replacing vLLM or TensorRT-LLM with a compiled engine; An open-source engine teams can run on their own hardware.

Is Baseten or Luminal cheaper?

Baseten: H100 about $6.50/hr dedicated. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.