# Modal vs Moonshot AI

> Moonshot AI serves Kimi K3, a 2.8T open model, at $3 in and $15 out. Modal rents per-second GPUs for models you bring, but K3 needs a 64+ accelerator cluster, beyond Modal's 8-GPU containers.

Canonical: https://www.subconscious.dev/compare/modal-vs-moonshot-ai · By The Subconscious Team · Updated September 30, 2026

## How they compare

Moonshot AI is the lab behind Kimi K3, a 2.8 trillion parameter mixture-of-experts model with 1M context and native vision. Vals AI scored it 93.4% on SWE-bench Verified, fourth overall. The hosted API charges $3 in and $15 out per million, with cached input at $0.30, and runs through an OpenAI-compatible endpoint, Kimi Code and OpenRouter. Modal sells serverless GPU containers billed per second, with up to 8 GPUs per container across T4 through B300. It has no models of its own.

Kimi K3 is a poor fit for self-hosting on Modal. Moonshot says self-hosting takes a 64+ accelerator cluster, far more than one 8-GPU container. For K3, the API is the practical path, with the caveats of about 33 tokens per second, verbose always-on thinking and the custom license. Modal fits the smaller work around it: embeddings, reranking, a fine-tuned small model, or agent sandboxes. The cheaper Kimi K2.6 stays available on Moonshot's API at $0.95 in and $4 out.

## What each one does

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

### Moonshot AI

Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.

## Which is best, and when

### Choose Modal for

- Custom small models, embeddings and reranking on demand.
- Agent sandboxes and batch jobs billed by the second.
- Fine-tunes you control end to end.

### Choose Moonshot AI for

- Near-frontier open-model coding on huge repositories.
- 1M-context document and visual agent work.
- Using Kimi K3 without running a 64+ accelerator cluster.

## At a glance

| Attribute | Modal | Moonshot AI |
|---|---|---|
| Model access | Bring your own weights | Open weights, custom license |
| Flagship models | None hosted | Kimi K3, Kimi K2.6 |
| Speed | ~1s container boot | ~33 tok/s on Kimi K3 |
| Price | Per second; H100 $3.95/hr list | $3 in, $15 out (Kimi K3) |
| Customization | Run any training code | Open weights to fine-tune |
| Deployment | Serverless GPU containers | API, Kimi Code, OpenRouter |
| Long context | Depends on the model you deploy | 1M |

## FAQ

### What is the difference between Modal and Moonshot AI?

Moonshot AI serves Kimi K3, a 2.8T open model, at $3 in and $15 out. Modal rents per-second GPUs for models you bring, but K3 needs a 64+ accelerator cluster, beyond Modal's 8-GPU containers.

### When should I choose Modal over Moonshot AI?

Custom small models, embeddings and reranking on demand; Agent sandboxes and batch jobs billed by the second; Fine-tunes you control end to end.

### When should I choose Moonshot AI over Modal?

Near-frontier open-model coding on huge repositories; 1M-context document and visual agent work; Using Kimi K3 without running a 64+ accelerator cluster.

### Is Modal or Moonshot AI cheaper?

Modal: Per second; H100 $3.95/hr list. Moonshot AI: $3 in, $15 out (Kimi K3). The cheaper choice depends on the model and workload.

### Which has more context, Modal or Moonshot AI?

Modal: Depends on the model you deploy. Moonshot AI: 1M.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md), [Subconscious vs Moonshot AI](https://www.subconscious.dev/compare/subconscious-vs-moonshot-ai.md).

Full profiles: [Modal](https://www.subconscious.dev/providers/modal.md), [Moonshot AI](https://www.subconscious.dev/providers/moonshot-ai.md).
