vs

xAI vs Moonshot AI

Grok is a closed, fast, cheap-output model with live X data. Kimi K3 is the most capable open-weight model, near frontier coding scores but slow and verbose.

By The Subconscious Team · Updated

xAI vs Moonshot AI: key differences

The split is speed and openness against raw coding strength. Moonshot's Kimi K3 is a 2.8 trillion parameter open-weight model with native vision and 1M context. Vals AI scored it 93.4% on SWE-bench Verified, fourth overall, and it costs $3 in and $15 out per million on Moonshot's API. It also always thinks and runs around 33 tokens per second. xAI's docs pitch Grok 4.20 on speed, strict prompt adherence and a low hallucination rate, and Grok output tokens are cheap: $6 per million on Grok 4.6 and $2.50 on Grok 4.20.

For coding agents on huge repositories, Kimi K3's scores and 1M context are the draw, and its open weights let a team self-host, though that takes a 64+ accelerator cluster and a custom license adds terms above $20M in hosting revenue. xAI answers with grok-build, a dedicated coding model, and 1M context on Grok 4.20, but its bill doubles past 200K prompt tokens. Moonshot's capacity has also been tight, with new API subscriptions paused after launch. Grok is the easier default for fast tool-calling and fresh-data agents.

What xAI and Moonshot AI do

xAI

xAI sells the Grok models through its own API. Grok 4.6 is the current flagship and xAI tells developers to use it for everything outside audio, image and video, code included. It has a 500K context window and costs $2 in and $6 out per million tokens under 200K prompt tokens. Older Grok 4.20 and 4.3 models keep a 1M window at $1.25 in and $2.50 out, which is aggressive for that capability class.

Example models: Grok 4.6, Grok 4.20

Full xAI profile

Moonshot AI

Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.

Example models: Kimi K3, Kimi K2.6

Full Moonshot AI profile

Should you choose xAI or Moonshot AI?

xAI

Choose xAI for

  • Fast, cheap-output reasoning and tool calls
  • Real-time agents that pull from X Search
  • Teams that do not want to run open weights

Moonshot AI

Choose Moonshot AI for

  • Repo-scale coding agents that want top open-weight scores
  • Visual and document-heavy research at 1M context
  • Teams that want the option to self-host

xAI vs Moonshot AI at a glance

AttributexAIMoonshot AI
Model accessClosedOpen weights, custom license
Flagship modelsGrok 4.6, Grok 4.20, grok-buildKimi K3, Kimi K2.6
Speed~54 tok/s on Grok 4.6~33 tok/s on Kimi K3
Price$2 in, $6 out (Grok 4.6); 2x past 200K$3 in, $15 out (Kimi K3)
CustomizationUnknownOpen weights to fine-tune
DeploymentFirst-party APIAPI, Kimi Code, OpenRouter
Long context500K (4.6), 1M (4.20, 4.3)1M

Frequently asked questions

What is the difference between xAI and Moonshot AI?

Grok is a closed, fast, cheap-output model with live X data. Kimi K3 is the most capable open-weight model, near frontier coding scores but slow and verbose.

When should I choose xAI over Moonshot AI?

Fast, cheap-output reasoning and tool calls; Real-time agents that pull from X Search; Teams that do not want to run open weights.

When should I choose Moonshot AI over xAI?

Repo-scale coding agents that want top open-weight scores; Visual and document-heavy research at 1M context; Teams that want the option to self-host.

Is xAI or Moonshot AI cheaper?

xAI: $2 in, $6 out (Grok 4.6); 2x past 200K. Moonshot AI: $3 in, $15 out (Kimi K3). The cheaper choice depends on the model and workload.

Which has more context, xAI or Moonshot AI?

xAI: 500K (4.6), 1M (4.20, 4.3). Moonshot AI: 1M.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.