Long-running agents deserve better inference.
vs

Moonshot AI vs Luminal

Moonshot makes Kimi K3, a top open-weight model that is slow and hard to self-host. Luminal compiles open models for more GPU throughput.

By The Subconscious Team · Updated

Moonshot AI vs Luminal: key differences

Moonshot's Kimi K3 is among the most capable open-weight models, with 1M context, served on its API at $3 in and $15 out, but it runs at about 33 tokens per second and self-hosting takes a 64+ accelerator cluster. Luminal builds an engine, not a model: a compiler that lowers models to primitive ops and emits fused kernels ahead of time.

Moonshot is the pick for model quality. Luminal is a pick for serving efficiency, with a reported 36K tokens per second on GPT-OSS 120B across 8 H100s. Luminal has not published Kimi results, and a model of K3's size is a much harder compile target, so teams should test before assuming the gains carry over.

What Moonshot AI and Luminal do

Moonshot AI

Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.

Example models: Kimi K3, Kimi K2.6

Full Moonshot AI profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Moonshot AI or Luminal?

Moonshot AI

Choose Moonshot AI for

  • Frontier open-weight quality
  • 1M context
  • Kimi Code for agentic coding

Luminal

Choose Luminal for

  • Maximum throughput per GPU on a self-chosen model
  • On-prem deployments with custom kernel work and SLAs
  • An open-source engine teams can run on their own hardware

Moonshot AI vs Luminal at a glance

AttributeMoonshot AILuminal
Model accessOpen weights, custom licenseBring your own weights
Flagship modelsKimi K3, Kimi K2.6No public catalog
Speed~33 tok/s on Kimi K336K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price$3 in, $15 out (Kimi K3)Pay per use; rates not published
CustomizationOpen weights to fine-tuneCompiles any PyTorch or HF model
DeploymentAPI, Kimi Code, OpenRouterServerless (early access), on-prem license
Long context1MUnknown

Frequently asked questions

What is the difference between Moonshot AI and Luminal?

Moonshot makes Kimi K3, a top open-weight model that is slow and hard to self-host. Luminal compiles open models for more GPU throughput.

When should I choose Moonshot AI over Luminal?

Frontier open-weight quality; 1M context; Kimi Code for agentic coding.

When should I choose Luminal over Moonshot AI?

Maximum throughput per GPU on a self-chosen model; On-prem deployments with custom kernel work and SLAs; An open-source engine teams can run on their own hardware.

Is Moonshot AI or Luminal cheaper?

Moonshot AI: $3 in, $15 out (Kimi K3). Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.