# Moonshot AI vs Thinking Machines

> Moonshot builds Kimi K3, the most capable open-weight model. Thinking Machines supports post-training Kimi K2.6 through Tinker and ships its own Inkling models.

Canonical: https://www.subconscious.dev/compare/moonshot-ai-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Kimi K3 is a 2.8 trillion parameter MoE with native vision and 1M context, and Vals AI scored it 93.4% on SWE-bench Verified, fourth overall. Moonshot's API charges $3 in and $15 out, with cached input at $0.30, and Kimi K2.6 costs $0.95 in and $4 out. Self-hosting K3 takes a 64+ accelerator cluster, and the custom license adds a commercial agreement above $20M in hosting revenue. Thinking Machines' Inkling is smaller, 975B total with 41B active, under a plain Apache 2.0 license, with text, image and audio input, 1M context, and pricing of $1.00 in and $4.05 out on a beta serverless API.

The two also connect. Tinker lists Kimi K2.6 as a trainable base, so a team can run LoRA SFT or RL on it without assembling a cluster, which is hard to do in-house. Moonshot's serving is more mature, reaching developers through an OpenAI-compatible API, Kimi Code, OpenRouter and Cloudflare Workers AI, though K3 is slow at around 33 tokens per second and demand briefly paused new API subscriptions in July. Thinking Machines only serves Inkling and keeps checkpoint sampling to low internal traffic. Moonshot wins for top open coding quality; Thinking Machines wins on license simplicity and training control.

## What each one does

### Moonshot AI

Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Moonshot AI for

- Top open-weight coding scores with Kimi K3
- Long-horizon coding on huge repositories
- Kimi Code in the terminal

### Choose Thinking Machines for

- LoRA post-training of Kimi K2.6 without a cluster
- Apache 2.0 weights with no revenue thresholds
- Audio and image input on Inkling

## At a glance

| Attribute | Moonshot AI | Thinking Machines |
|---|---|---|
| Model access | Open weights, custom license | Open weights |
| Flagship models | Kimi K3, Kimi K2.6 | Inkling, Inkling-Small |
| Speed | ~33 tok/s on Kimi K3 | - |
| Price | $3 in, $15 out (Kimi K3) | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Open weights to fine-tune | LoRA SFT and RL via Tinker |
| Deployment | API, Kimi Code, OpenRouter | Training API, beta serverless (Inkling only) |
| Long context | 1M | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Moonshot AI and Thinking Machines?

Moonshot builds Kimi K3, the most capable open-weight model. Thinking Machines supports post-training Kimi K2.6 through Tinker and ships its own Inkling models.

### When should I choose Moonshot AI over Thinking Machines?

Top open-weight coding scores with Kimi K3; Long-horizon coding on huge repositories; Kimi Code in the terminal.

### When should I choose Thinking Machines over Moonshot AI?

LoRA post-training of Kimi K2.6 without a cluster; Apache 2.0 weights with no revenue thresholds; Audio and image input on Inkling.

### Is Moonshot AI or Thinking Machines cheaper?

Moonshot AI: $3 in, $15 out (Kimi K3). Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Moonshot AI or Thinking Machines?

Moonshot AI: 1M. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Moonshot AI](https://www.subconscious.dev/compare/subconscious-vs-moonshot-ai.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Moonshot AI](https://www.subconscious.dev/providers/moonshot-ai.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
