Baseten vs Moonshot AI
Moonshot built Kimi K3 and sells it at $3 in and $15 out. Baseten also serves Kimi K3, alongside 12 other open models, on GPUs tuned for fast first tokens.
By The Subconscious Team · Updated
Baseten vs Moonshot AI: key differences
Moonshot AI is the lab behind Kimi K3, a 2.8 trillion parameter mixture-of-experts model with a 1M token context. Vals AI placed it fourth on SWE-bench Verified behind only closed frontier models. Moonshot's hosted API charges $3 in and $15 out per million, with cached input at $0.30. Its weak spot is capacity and speed: K3 runs around 33 tokens per second, and demand overran Moonshot's GPUs badly enough that new API subscriptions paused on July 19 before reopening in batches. Baseten lists Kimi K3 in its curated catalog, so a team can reach the same weights on a host whose stack is built for low time to first token.
Beyond Kimi, the difference is scope. Moonshot offers one family plus Kimi Code for the terminal. Baseten covers DeepSeek V4, GLM 5.2 and gpt-oss 120B too, and it handles compliance with HIPAA and data residency. Self-hosting K3 takes a 64+ accelerator cluster, and the license adds a commercial agreement above $20M in hosting revenue, so most teams will call an API. Moonshot also has the cheaper Kimi K2.6 at $0.95 in and $4 out.
What Baseten and Moonshot AI do
Baseten
Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.
Example models: GLM 5.2, gpt-oss 120B
Full Baseten profileMoonshot AI
Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.
Example models: Kimi K3, Kimi K2.6
Full Moonshot AI profileShould you choose Baseten or Moonshot AI?
Baseten
Choose Baseten for
- Kimi K3 with HIPAA or data residency requirements
- Switching between Kimi, DeepSeek and GLM on one key
- Teams burned by Moonshot capacity limits
Moonshot AI
Choose Moonshot AI for
- Direct first-party access to Kimi K3 and K2.6
- Using Kimi Code in the terminal
- Repo-scale agents that rely on cached input at $0.30
Baseten vs Moonshot AI at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights, 13 curated | Open weights, custom license |
| Flagship models | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | Kimi K3, Kimi K2.6 |
| Speed | 0.49s TTFT, lowest measured | ~33 tok/s on Kimi K3 |
| Price | H100 about $6.50/hr dedicated | $3 in, $15 out (Kimi K3) |
| Customization | Deploy any model with Truss | Open weights to fine-tune |
| Deployment | Model APIs, dedicated, self-host | API, Kimi Code, OpenRouter |
| Long context | Varies by model | 1M |
Frequently asked questions
What is the difference between Baseten and Moonshot AI?
Moonshot built Kimi K3 and sells it at $3 in and $15 out. Baseten also serves Kimi K3, alongside 12 other open models, on GPUs tuned for fast first tokens.
When should I choose Baseten over Moonshot AI?
Kimi K3 with HIPAA or data residency requirements; Switching between Kimi, DeepSeek and GLM on one key; Teams burned by Moonshot capacity limits.
When should I choose Moonshot AI over Baseten?
Direct first-party access to Kimi K3 and K2.6; Using Kimi Code in the terminal; Repo-scale agents that rely on cached input at $0.30.
Is Baseten or Moonshot AI cheaper?
Baseten: H100 about $6.50/hr dedicated. Moonshot AI: $3 in, $15 out (Kimi K3). The cheaper choice depends on the model and workload.
Which has more context, Baseten or Moonshot AI?
Baseten: Varies by model. Moonshot AI: 1M.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.