# Subconscious vs Moonshot AI

> Kimi K3 posts top open-model coding scores but runs around 33 tokens per second. Subconscious keeps long agents moving with 2x faster task completion and processed-token billing.

Canonical: https://www.subconscious.dev/compare/subconscious-vs-moonshot-ai · By The Subconscious Team · Updated September 30, 2026

## How they compare

Kimi K3 is the strongest open model in this guide. Vals AI scored it 93.4% on SWE-bench Verified, fourth overall behind closed models, and it offers a 1M context aimed at repo-scale agents. It also always thinks and runs around 33 tokens per second, at $3 in and $15 out on Moonshot's API. On an hour-long coding run, that speed and verbosity compound. Subconscious takes a different route. Its managed API runs GLM 5.3 and DeepSeek V4.1 Flash through a runtime that prunes the KV cache, delivers 2x faster task completion and a 5M+ effective context window, and bills only processed tokens.

Capacity and licensing also separate them. Demand for K3 overran Moonshot's GPUs within days of launch, and new API subscriptions paused before reopening in batches. Its custom license adds a commercial agreement above $20M in hosting revenue, and self-hosting K3 takes a 64+ accelerator cluster. When a task needs the best open coding score and can afford to wait, Moonshot is the pick. When a long agent has to move quickly and cheaply through millions of tokens, Subconscious fits better, and its dedicated deployments can run nearly any open model for teams with a specific checkpoint in mind.

## What each one does

### Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### Moonshot AI

Moonshot AI is the Beijing lab behind the Kimi models. Its flagship Kimi K3 launched July 16, 2026 as a 2.8 trillion parameter mixture-of-experts model that activates 16 of 896 experts per token, with native vision and a 1M token context. It is the first open model in the 3T class, and full weights landed on Hugging Face on July 27. The hosted API costs $3 in and $15 out per million tokens, with cached input at $0.30, and it runs through an OpenAI-compatible endpoint, Kimi Code in the terminal, OpenRouter and Cloudflare Workers AI.

## Which is best, and when

### Choose Subconscious for

- Long runs where ~33 tokens per second would stall the loop
- Traces that run past 1M tokens
- Processed-token billing on hour-long coding sessions

### Choose Moonshot AI for

- The hardest coding tasks, where open-model quality matters most
- Visual and document-heavy agents needing native vision and 1M context
- Fine-tuning the most capable open weights

## At a glance

| Attribute | Subconscious | Moonshot AI |
|---|---|---|
| Model access | Open weights | Open weights, custom license |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | Kimi K3, Kimi K2.6 |
| Speed | 2x faster task completion | ~33 tok/s on Kimi K3 |
| Price | 50–80% lower cost; billed on processed tokens | $3 in, $15 out (Kimi K3) |
| Customization | Marathon post-trained variants | Open weights to fine-tune |
| Deployment | Managed API, dedicated, on-prem | API, Kimi Code, OpenRouter |
| Long context | 5M+ effective context | 1M |

## FAQ

### What is the difference between Subconscious and Moonshot AI?

Kimi K3 posts top open-model coding scores but runs around 33 tokens per second. Subconscious keeps long agents moving with 2x faster task completion and processed-token billing.

### When should I choose Subconscious over Moonshot AI?

Long runs where ~33 tokens per second would stall the loop; Traces that run past 1M tokens; Processed-token billing on hour-long coding sessions.

### When should I choose Moonshot AI over Subconscious?

The hardest coding tasks, where open-model quality matters most; Visual and document-heavy agents needing native vision and 1M context; Fine-tuning the most capable open weights.

### Is Subconscious or Moonshot AI cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Moonshot AI: $3 in, $15 out (Kimi K3). The cheaper choice depends on the model and workload.

### Which has more context, Subconscious or Moonshot AI?

Subconscious: 5M+ effective context. Moonshot AI: 1M.

## Try Subconscious

Subconscious speaks the OpenAI and Anthropic API formats. Base URL: https://api.subconscious.dev/v1. Docs: https://docs.subconscious.dev. Get an API key: https://platform.subconscious.dev/signin. Agent guide: https://www.subconscious.dev/agents.md.

Full profiles: [Subconscious](https://www.subconscious.dev/providers/subconscious.md), [Moonshot AI](https://www.subconscious.dev/providers/moonshot-ai.md).
