Long-running agents deserve better inference.
vs

Crusoe vs Luminal

Crusoe builds its own data centers and serves open models on a shared KV cache. Luminal compiles models into faster kernels for any GPU.

By The Subconscious Team · Updated

Crusoe vs Luminal: key differences

Crusoe is an energy-first AI cloud with serverless open models, LoRA fine-tuning, dedicated endpoints and raw GPUs, and a cluster-wide KV cache it says makes time to first token up to 9.9x faster than vLLM on prefix-heavy work. Luminal works at the kernel level: its compiler lowers a model to primitive ops and emits fused native code ahead of time.

Both measure against vLLM and both numbers come from the vendors. Crusoe's gains come from cache reuse across requests; Luminal's from faster compute per step, with 36K tokens per second reported on GPT-OSS 120B over 8 H100s. Crusoe brings capacity and a catalog. Luminal brings an engine you can also run on-prem.

What Crusoe and Luminal do

Crusoe

Crusoe started in 2018 turning wasted natural gas into power for computing and has since become a vertically integrated AI infrastructure company: it sources energy, builds data centers and rents GPUs through Crusoe Cloud. It designed and built the Abilene, Texas campus behind the OpenAI and Oracle Stargate project, planned at 1.2 GW, and in March 2026 announced an adjacent 900 MW campus for Microsoft. On September 17, 2026 it closed the first part of a $3.9B Series F at a $30.9B post-money valuation, and it reports over 6 GW of contracted capacity. Crusoe Cloud lists GB200 NVL72, B200 and AMD MI355X by quote, with H100 at $3.90 and H200 at $4.29 per GPU-hour on demand.

Example models: DeepSeek V4 Pro, GLM 5.3, Kimi K2.6

Full Crusoe profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Crusoe or Luminal?

Crusoe

Choose Crusoe for

  • Prefix-heavy workloads that reuse cached context
  • Capacity from an owned data center fleet
  • Serverless LoRA fine-tuning

Luminal

Choose Luminal for

  • On-prem deployments with custom kernel work and SLAs
  • Serving custom or fine-tuned architectures off any catalog
  • An open-source engine teams can run on their own hardware

Crusoe vs Luminal at a glance

AttributeCrusoeLuminal
Model accessOpen weightsBring your own weights
Flagship modelsDeepSeek V4, GLM 5.3, Kimi K2.6, Nemotron 3No public catalog
SpeedUp to 9.9x faster TTFT vs vLLM (vendor claim)36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price$0.05–$1.74 in, $0.20–$4.40 out per 1MPay per use; rates not published
CustomizationServerless LoRA fine-tuningCompiles any PyTorch or HF model
DeploymentServerless, self-serve and tailored dedicated, raw GPUsServerless (early access), on-prem license
Long contextVaries by model; cluster-wide KV cacheUnknown

Frequently asked questions

What is the difference between Crusoe and Luminal?

Crusoe builds its own data centers and serves open models on a shared KV cache. Luminal compiles models into faster kernels for any GPU.

When should I choose Crusoe over Luminal?

Prefix-heavy workloads that reuse cached context; Capacity from an owned data center fleet; Serverless LoRA fine-tuning.

When should I choose Luminal over Crusoe?

On-prem deployments with custom kernel work and SLAs; Serving custom or fine-tuned architectures off any catalog; An open-source engine teams can run on their own hardware.

Is Crusoe or Luminal cheaper?

Crusoe: $0.05–$1.74 in, $0.20–$4.40 out per 1M. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.