Long-running agents deserve better inference.
vs

Inference.net vs Luminal

Inference.net sells cheap batch on spare GPUs and distills traces into custom models. Luminal compiles models so each GPU does more.

By The Subconscious Team · Updated

Inference.net vs Luminal: key differences

Inference.net turns spare GPU capacity into cheap batch inference with 24-hour to 7-day windows, and helps teams distill production traces into smaller custom models. Luminal works on the engine: its compiler lowers a model into native kernels ahead of time, reporting 36K tokens per second on GPT-OSS 120B across 8 H100s.

The two attack cost differently. Inference.net cuts cost by waiting for idle capacity and by shrinking the model. Luminal cuts it by getting more out of each GPU, which also helps real-time serving. A team could distill a model with Inference.net and serve it through Luminal.

What Inference.net and Luminal do

Inference.net

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

Example models: open catalog models plus customer fine-tunes served on dedicated GPUs

Full Inference.net profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Inference.net or Luminal?

Inference.net

Choose Inference.net for

  • Cheap batch with flexible deadlines
  • Distilling traces into custom models
  • A gateway across model sources

Luminal

Choose Luminal for

  • Real-time serving of a custom model at high throughput
  • On-prem deployments with custom kernel work and SLAs
  • An open-source engine teams can run on their own hardware

Inference.net vs Luminal at a glance

AttributeInference.netLuminal
Model accessOpen, closed and customBring your own weights
Flagship modelsCustomer fine-tunesNo public catalog
SpeedBatch windows of 24h to 7 days36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
PriceDiscounted spare GPU capacityPay per use; rates not published
CustomizationDistill traces into custom modelsCompiles any PyTorch or HF model
DeploymentBatch API, gateway, dedicated GPUsServerless (early access), on-prem license
Long contextVaries by modelUnknown

Frequently asked questions

What is the difference between Inference.net and Luminal?

Inference.net sells cheap batch on spare GPUs and distills traces into custom models. Luminal compiles models so each GPU does more.

When should I choose Inference.net over Luminal?

Cheap batch with flexible deadlines; Distilling traces into custom models; A gateway across model sources.

When should I choose Luminal over Inference.net?

Real-time serving of a custom model at high throughput; On-prem deployments with custom kernel work and SLAs; An open-source engine teams can run on their own hardware.

Is Inference.net or Luminal cheaper?

Inference.net: Discounted spare GPU capacity. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.