Long-running agents deserve better inference.
vs

xAI vs Luminal

xAI sells closed Grok models with live X data. Luminal compiles open models into native GPU code for teams that run their own weights.

By The Subconscious Team · Updated

xAI vs Luminal: key differences

xAI offers Grok 4.6, Grok 4.20 and grok-build through its own API, with live data from X and cheap output tokens, though prompts past 200K tokens bill at double. Luminal sells no model. Its compiler takes an open model you supply and turns it into fused GPU kernels ahead of time, served serverless in early access or on-prem.

Choose xAI for a closed model with real-time social data. Choose Luminal when you have settled on open weights and want more throughput from your hardware; it reports GPT-OSS 120B at 36K tokens per second over 8 H100s. Luminal has no published pricing and no frontier model of its own.

What xAI and Luminal do

xAI

xAI sells the Grok models through its own API. Grok 4.6 is the current flagship and xAI tells developers to use it for everything outside audio, image and video, code included. It has a 500K context window and costs $2 in and $6 out per million tokens under 200K prompt tokens. Older Grok 4.20 and 4.3 models keep a 1M window at $1.25 in and $2.50 out, which is aggressive for that capability class.

Example models: Grok 4.6, Grok 4.20

Full xAI profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose xAI or Luminal?

xAI

Choose xAI for

  • Live X data inside model answers
  • Cheap output tokens on a closed model
  • grok-build for coding

Luminal

Choose Luminal for

  • Compiling a custom model into fast native GPU code
  • An open-source engine teams can run on their own hardware
  • On-prem deployments with custom kernel work and SLAs

xAI vs Luminal at a glance

AttributexAILuminal
Model accessClosedBring your own weights
Flagship modelsGrok 4.6, Grok 4.20, grok-buildNo public catalog
Speed~54 tok/s on Grok 4.636K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price$2 in, $6 out (Grok 4.6); 2x past 200KPay per use; rates not published
CustomizationUnknownCompiles any PyTorch or HF model
DeploymentFirst-party APIServerless (early access), on-prem license
Long context500K (4.6), 1M (4.20, 4.3)Unknown

Frequently asked questions

What is the difference between xAI and Luminal?

xAI sells closed Grok models with live X data. Luminal compiles open models into native GPU code for teams that run their own weights.

When should I choose xAI over Luminal?

Live X data inside model answers; Cheap output tokens on a closed model; grok-build for coding.

When should I choose Luminal over xAI?

Compiling a custom model into fast native GPU code; An open-source engine teams can run on their own hardware; On-prem deployments with custom kernel work and SLAs.

Is xAI or Luminal cheaper?

xAI: $2 in, $6 out (Grok 4.6); 2x past 200K. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.