Modal vs Moonshot AI
Moonshot AI serves Kimi K3, a 2.8T open model, at $3 in and $15 out. Modal rents per-second GPUs for models you bring, but K3 needs a 64+ accelerator cluster, beyond Modal's 8-GPU containers.
By The Subconscious Team · Updated
Modal vs Moonshot AI: key differences
Moonshot AI is the lab behind Kimi K3, a 2.8 trillion parameter mixture-of-experts model with 1M context and native vision. Vals AI scored it 93.4% on SWE-bench Verified, fourth overall. The hosted API charges $3 in and $15 out per million, with cached input at $0.30, and runs through an OpenAI-compatible endpoint, Kimi Code and OpenRouter. Modal sells serverless GPU containers billed per second, with up to 8 GPUs per container across T4 through B300. It has no models of its own.
Kimi K3 is a poor fit for self-hosting on Modal. Moonshot says self-hosting takes a 64+ accelerator cluster, far more than one 8-GPU container. For K3, the API is the practical path, with the caveats of about 33 tokens per second, verbose always-on thinking and the custom license. Modal fits the smaller work around it: embeddings, reranking, a fine-tuned small model, or agent sandboxes. The cheaper Kimi K2.6 stays available on Moonshot's API at $0.95 in and $4 out.
What Modal and Moonshot AI do
Modal
Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.
Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper
Full Modal profileMoonshot AI
Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.
Example models: Kimi K3, Kimi K2.6
Full Moonshot AI profileShould you choose Modal or Moonshot AI?
Modal
Choose Modal for
- Custom small models, embeddings and reranking on demand.
- Agent sandboxes and batch jobs billed by the second.
- Fine-tunes you control end to end.
Moonshot AI
Choose Moonshot AI for
- Near-frontier open-model coding on huge repositories.
- 1M-context document and visual agent work.
- Using Kimi K3 without running a 64+ accelerator cluster.
Modal vs Moonshot AI at a glance
| Attribute | ||
|---|---|---|
| Model access | Bring your own weights | Open weights, custom license |
| Flagship models | None hosted | Kimi K3, Kimi K2.6 |
| Speed | ~1s container boot | ~33 tok/s on Kimi K3 |
| Price | Per second; H100 $3.95/hr list | $3 in, $15 out (Kimi K3) |
| Customization | Run any training code | Open weights to fine-tune |
| Deployment | Serverless GPU containers | API, Kimi Code, OpenRouter |
| Long context | Depends on the model you deploy | 1M |
Frequently asked questions
What is the difference between Modal and Moonshot AI?
Moonshot AI serves Kimi K3, a 2.8T open model, at $3 in and $15 out. Modal rents per-second GPUs for models you bring, but K3 needs a 64+ accelerator cluster, beyond Modal's 8-GPU containers.
When should I choose Modal over Moonshot AI?
Custom small models, embeddings and reranking on demand; Agent sandboxes and batch jobs billed by the second; Fine-tunes you control end to end.
When should I choose Moonshot AI over Modal?
Near-frontier open-model coding on huge repositories; 1M-context document and visual agent work; Using Kimi K3 without running a 64+ accelerator cluster.
Is Modal or Moonshot AI cheaper?
Modal: Per second; H100 $3.95/hr list. Moonshot AI: $3 in, $15 out (Kimi K3). The cheaper choice depends on the model and workload.
Which has more context, Modal or Moonshot AI?
Modal: Depends on the model you deploy. Moonshot AI: 1M.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.