vs

Modal vs Moonshot AI

Moonshot AI serves Kimi K3, a 2.8T open model, at $3 in and $15 out. Modal rents per-second GPUs for models you bring, but K3 needs a 64+ accelerator cluster, beyond Modal's 8-GPU containers.

By The Subconscious Team · Updated

Modal vs Moonshot AI: key differences

Moonshot AI is the lab behind Kimi K3, a 2.8 trillion parameter mixture-of-experts model with 1M context and native vision. Vals AI scored it 93.4% on SWE-bench Verified, fourth overall. The hosted API charges $3 in and $15 out per million, with cached input at $0.30, and runs through an OpenAI-compatible endpoint, Kimi Code and OpenRouter. Modal sells serverless GPU containers billed per second, with up to 8 GPUs per container across T4 through B300. It has no models of its own.

Kimi K3 is a poor fit for self-hosting on Modal. Moonshot says self-hosting takes a 64+ accelerator cluster, far more than one 8-GPU container. For K3, the API is the practical path, with the caveats of about 33 tokens per second, verbose always-on thinking and the custom license. Modal fits the smaller work around it: embeddings, reranking, a fine-tuned small model, or agent sandboxes. The cheaper Kimi K2.6 stays available on Moonshot's API at $0.95 in and $4 out.

What Modal and Moonshot AI do

Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper

Full Modal profile

Moonshot AI

Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.

Example models: Kimi K3, Kimi K2.6

Full Moonshot AI profile

Should you choose Modal or Moonshot AI?

Modal

Choose Modal for

  • Custom small models, embeddings and reranking on demand.
  • Agent sandboxes and batch jobs billed by the second.
  • Fine-tunes you control end to end.

Moonshot AI

Choose Moonshot AI for

  • Near-frontier open-model coding on huge repositories.
  • 1M-context document and visual agent work.
  • Using Kimi K3 without running a 64+ accelerator cluster.

Modal vs Moonshot AI at a glance

AttributeModalMoonshot AI
Model accessBring your own weightsOpen weights, custom license
Flagship modelsNone hostedKimi K3, Kimi K2.6
Speed~1s container boot~33 tok/s on Kimi K3
PricePer second; H100 $3.95/hr list$3 in, $15 out (Kimi K3)
CustomizationRun any training codeOpen weights to fine-tune
DeploymentServerless GPU containersAPI, Kimi Code, OpenRouter
Long contextDepends on the model you deploy1M

Frequently asked questions

What is the difference between Modal and Moonshot AI?

Moonshot AI serves Kimi K3, a 2.8T open model, at $3 in and $15 out. Modal rents per-second GPUs for models you bring, but K3 needs a 64+ accelerator cluster, beyond Modal's 8-GPU containers.

When should I choose Modal over Moonshot AI?

Custom small models, embeddings and reranking on demand; Agent sandboxes and batch jobs billed by the second; Fine-tunes you control end to end.

When should I choose Moonshot AI over Modal?

Near-frontier open-model coding on huge repositories; 1M-context document and visual agent work; Using Kimi K3 without running a 64+ accelerator cluster.

Is Modal or Moonshot AI cheaper?

Modal: Per second; H100 $3.95/hr list. Moonshot AI: $3 in, $15 out (Kimi K3). The cheaper choice depends on the model and workload.

Which has more context, Modal or Moonshot AI?

Modal: Depends on the model you deploy. Moonshot AI: 1M.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.