RunInfra vs Luminal
RunInfra uses an agent to benchmark and quantize deployments; Luminal compiles models into native kernels. Both chase speed on open models.
By The Subconscious Team · Updated
RunInfra vs Luminal: key differences
RunInfra's agent takes a plain-English request, benchmarks models across GPUs from L4 to B200, tries AWQ, GPTQ and FP8 variants, applies its Forge kernels and ships an endpoint that scales to zero. It also sells coding plans from $10 a month. Luminal compiles a model into fused native kernels ahead of time, reporting 36K tokens per second on GPT-OSS 120B across 8 H100s.
RunInfra is the friendlier product for small teams: published plans, custom uploads up to 50 GB and voice pipelines. Luminal goes deeper on the engine and offers an open-source compiler and on-prem license, but its cloud is early access with no public prices. Both are young, so benchmark on your own traffic.
What RunInfra and Luminal do
RunInfra
RunInfra pitches open models built for agents, with two ways in. Its hosted Model APIs serve a small curated library, including Nemotron 3.5 Lightning 30B, Qwen 3.8 27B and Ornith 1.5 35B, behind one key that works with both the OpenAI and Anthropic SDKs. Cached context bills at a discount. Coding plans start at $10 a month with limits that reset every five hours and every week, and they plug into Claude Code, Codex, OpenCode, Cline, Aider and dozens of other agent CLIs.
Example models: Nemotron 3.5 Lightning 30B, Qwen 3.8 27B
Full RunInfra profileLuminal
Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.
Example models: GPT-OSS 120B, Llama 3 8B
Full Luminal profileShould you choose RunInfra or Luminal?
RunInfra
Choose RunInfra for
- Cheap coding plans in agent CLIs
- Automatic quantization toward a latency target
- Voice pipelines
Luminal
Choose Luminal for
- An open-source engine teams can run on their own hardware
- On-prem deployments with custom kernel work and SLAs
- Maximum throughput per GPU on a self-chosen model
RunInfra vs Luminal at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Bring your own weights |
| Flagship models | Nemotron 3.5 Lightning 30B, Qwen 3.8 27B | No public catalog |
| Speed | Cold starts under 2s | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | Coding plans from $10 a month | Pay per use; rates not published |
| Customization | Uploads up to 50 GB; auto-quantization | Compiles any PyTorch or HF model |
| Deployment | Model APIs, agent-built endpoints | Serverless (early access), on-prem license |
| Long context | Varies by model | Unknown |
Frequently asked questions
What is the difference between RunInfra and Luminal?
RunInfra uses an agent to benchmark and quantize deployments; Luminal compiles models into native kernels. Both chase speed on open models.
When should I choose RunInfra over Luminal?
Cheap coding plans in agent CLIs; Automatic quantization toward a latency target; Voice pipelines.
When should I choose Luminal over RunInfra?
An open-source engine teams can run on their own hardware; On-prem deployments with custom kernel work and SLAs; Maximum throughput per GPU on a self-chosen model.
Is RunInfra or Luminal cheaper?
RunInfra: Coding plans from $10 a month. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.