Moonshot AI vs Thinking Machines
Moonshot builds Kimi K3, the most capable open-weight model. Thinking Machines supports post-training Kimi K2.6 through Tinker and ships its own Inkling models.
By The Subconscious Team · Updated
Moonshot AI vs Thinking Machines: key differences
Kimi K3 is a 2.8 trillion parameter MoE with native vision and 1M context, and Vals AI scored it 93.4% on SWE-bench Verified, fourth overall. Moonshot's API charges $3 in and $15 out, with cached input at $0.30, and Kimi K2.6 costs $0.95 in and $4 out. Self-hosting K3 takes a 64+ accelerator cluster, and the custom license adds a commercial agreement above $20M in hosting revenue. Thinking Machines' Inkling is smaller, 975B total with 41B active, under a plain Apache 2.0 license, with text, image and audio input, 1M context, and pricing of $1.00 in and $4.05 out on a beta serverless API.
The two also connect. Tinker lists Kimi K2.6 as a trainable base, so a team can run LoRA SFT or RL on it without assembling a cluster, which is hard to do in-house. Moonshot's serving is more mature, reaching developers through an OpenAI-compatible API, Kimi Code, OpenRouter and Cloudflare Workers AI, though K3 is slow at around 33 tokens per second and demand briefly paused new API subscriptions in July. Thinking Machines only serves Inkling and keeps checkpoint sampling to low internal traffic. Moonshot wins for top open coding quality; Thinking Machines wins on license simplicity and training control.
What Moonshot AI and Thinking Machines do
Moonshot AI
Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.
Example models: Kimi K3, Kimi K2.6
Full Moonshot AI profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Moonshot AI or Thinking Machines?
Moonshot AI
Choose Moonshot AI for
- Top open-weight coding scores with Kimi K3
- Long-horizon coding on huge repositories
- Kimi Code in the terminal
Thinking Machines
Choose Thinking Machines for
- LoRA post-training of Kimi K2.6 without a cluster
- Apache 2.0 weights with no revenue thresholds
- Audio and image input on Inkling
Moonshot AI vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights, custom license | Open weights |
| Flagship models | Kimi K3, Kimi K2.6 | Inkling, Inkling-Small |
| Speed | ~33 tok/s on Kimi K3 | Unknown |
| Price | $3 in, $15 out (Kimi K3) | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Open weights to fine-tune | LoRA SFT and RL via Tinker |
| Deployment | API, Kimi Code, OpenRouter | Training API, beta serverless (Inkling only) |
| Long context | 1M | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Moonshot AI and Thinking Machines?
Moonshot builds Kimi K3, the most capable open-weight model. Thinking Machines supports post-training Kimi K2.6 through Tinker and ships its own Inkling models.
When should I choose Moonshot AI over Thinking Machines?
Top open-weight coding scores with Kimi K3; Long-horizon coding on huge repositories; Kimi Code in the terminal.
When should I choose Thinking Machines over Moonshot AI?
LoRA post-training of Kimi K2.6 without a cluster; Apache 2.0 weights with no revenue thresholds; Audio and image input on Inkling.
Is Moonshot AI or Thinking Machines cheaper?
Moonshot AI: $3 in, $15 out (Kimi K3). Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Moonshot AI or Thinking Machines?
Moonshot AI: 1M. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Moonshot AI
OpenAI vs Moonshot AI
Anthropic vs Moonshot AI
Google Vertex AI vs Moonshot AI
Amazon Bedrock vs Moonshot AI
Together AI vs Moonshot AI
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.