# Inference.net vs Luminal

> Inference.net sells cheap batch on spare GPUs and distills traces into custom models. Luminal compiles models so each GPU does more.

Canonical: https://www.subconscious.dev/compare/inference-net-vs-luminal · By The Subconscious Team · Updated September 30, 2026

## How they compare

Inference.net turns spare GPU capacity into cheap batch inference with 24-hour to 7-day windows, and helps teams distill production traces into smaller custom models. Luminal works on the engine: its compiler lowers a model into native kernels ahead of time, reporting 36K tokens per second on GPT-OSS 120B across 8 H100s.

The two attack cost differently. Inference.net cuts cost by waiting for idle capacity and by shrinking the model. Luminal cuts it by getting more out of each GPU, which also helps real-time serving. A team could distill a model with Inference.net and serve it through Luminal.

## What each one does

### Inference.net

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

### Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

## Which is best, and when

### Choose Inference.net for

- Cheap batch with flexible deadlines
- Distilling traces into custom models
- A gateway across model sources

### Choose Luminal for

- Real-time serving of a custom model at high throughput
- On-prem deployments with custom kernel work and SLAs
- An open-source engine teams can run on their own hardware

## At a glance

| Attribute | Inference.net | Luminal |
|---|---|---|
| Model access | Open, closed and custom | Bring your own weights |
| Flagship models | Customer fine-tunes | No public catalog |
| Speed | Batch windows of 24h to 7 days | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | Discounted spare GPU capacity | Pay per use; rates not published |
| Customization | Distill traces into custom models | Compiles any PyTorch or HF model |
| Deployment | Batch API, gateway, dedicated GPUs | Serverless (early access), on-prem license |
| Long context | Varies by model | - |

## FAQ

### What is the difference between Inference.net and Luminal?

Inference.net sells cheap batch on spare GPUs and distills traces into custom models. Luminal compiles models so each GPU does more.

### When should I choose Inference.net over Luminal?

Cheap batch with flexible deadlines; Distilling traces into custom models; A gateway across model sources.

### When should I choose Luminal over Inference.net?

Real-time serving of a custom model at high throughput; On-prem deployments with custom kernel work and SLAs; An open-source engine teams can run on their own hardware.

### Is Inference.net or Luminal cheaper?

Inference.net: Discounted spare GPU capacity. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Inference.net](https://www.subconscious.dev/compare/subconscious-vs-inference-net.md), [Subconscious vs Luminal](https://www.subconscious.dev/compare/subconscious-vs-luminal.md).

Full profiles: [Inference.net](https://www.subconscious.dev/providers/inference-net.md), [Luminal](https://www.subconscious.dev/providers/luminal.md).
