Moonshot AI vs Luminal
Moonshot makes Kimi K3, a top open-weight model that is slow and hard to self-host. Luminal compiles open models for more GPU throughput.
By The Subconscious Team · Updated
Moonshot AI vs Luminal: key differences
Moonshot's Kimi K3 is among the most capable open-weight models, with 1M context, served on its API at $3 in and $15 out, but it runs at about 33 tokens per second and self-hosting takes a 64+ accelerator cluster. Luminal builds an engine, not a model: a compiler that lowers models to primitive ops and emits fused kernels ahead of time.
Moonshot is the pick for model quality. Luminal is a pick for serving efficiency, with a reported 36K tokens per second on GPT-OSS 120B across 8 H100s. Luminal has not published Kimi results, and a model of K3's size is a much harder compile target, so teams should test before assuming the gains carry over.
What Moonshot AI and Luminal do
Moonshot AI
Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.
Example models: Kimi K3, Kimi K2.6
Full Moonshot AI profileLuminal
Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.
Example models: GPT-OSS 120B, Llama 3 8B
Full Luminal profileShould you choose Moonshot AI or Luminal?
Luminal
Choose Luminal for
- Maximum throughput per GPU on a self-chosen model
- On-prem deployments with custom kernel work and SLAs
- An open-source engine teams can run on their own hardware
Moonshot AI vs Luminal at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights, custom license | Bring your own weights |
| Flagship models | Kimi K3, Kimi K2.6 | No public catalog |
| Speed | ~33 tok/s on Kimi K3 | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | $3 in, $15 out (Kimi K3) | Pay per use; rates not published |
| Customization | Open weights to fine-tune | Compiles any PyTorch or HF model |
| Deployment | API, Kimi Code, OpenRouter | Serverless (early access), on-prem license |
| Long context | 1M | Unknown |
Frequently asked questions
What is the difference between Moonshot AI and Luminal?
Moonshot makes Kimi K3, a top open-weight model that is slow and hard to self-host. Luminal compiles open models for more GPU throughput.
When should I choose Moonshot AI over Luminal?
Frontier open-weight quality; 1M context; Kimi Code for agentic coding.
When should I choose Luminal over Moonshot AI?
Maximum throughput per GPU on a self-chosen model; On-prem deployments with custom kernel work and SLAs; An open-source engine teams can run on their own hardware.
Is Moonshot AI or Luminal cheaper?
Moonshot AI: $3 in, $15 out (Kimi K3). Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.