# Subconscious vs Crusoe

> Both serve open models like GLM 5.3 and both target agents that resend long context. Crusoe reuses KV cache across a cluster; Subconscious prunes it and bills processed tokens.

Canonical: https://www.subconscious.dev/compare/subconscious-vs-crusoe · By The Subconscious Team · Updated September 30, 2026

## How they compare

The overlap here is unusually direct. Both serve open models behind OpenAI-compatible APIs, GLM 5.3 sits on both menus, and both built their pitch around agents that resend long context. The mechanisms differ. Crusoe's MemoryAlloy shares a KV cache across the cluster with cache-aware routing, so a prefix computed on one node is reused on another, and cached input bills well below list. Crusoe claims up to 9.9x faster time to first token versus vLLM on prefix-heavy work. Subconscious prunes the KV cache and preserves suffix state instead of rereading a growing context. Against open models on standard inference it delivers 2x faster task completion, a 5M+ effective context window and 50–80% lower cost, billed on tokens processed after compression rather than tokens sent.

Crusoe wins on breadth of infrastructure. One account covers serverless tokens from $0.05 in, self-serve dedicated endpoints billed per GPU-hour, tailored deployments with SLAs, serverless LoRA fine-tuning and raw GPU clusters on Kubernetes or Slurm, including GB200 and B200 capacity by quote. Its serverless catalog is small but still wider than the Subconscious managed API, adding Kimi, Gemma, gpt-oss and Nemotron. Subconscious counters with Marathon post-trained variants co-designed with its runtime, on-prem deployment, no prompt logging, and direct plugs into Claude Code, Codex and Cursor. For short prefix-heavy chat, Crusoe's cache reuse is a good fit. Once a single agent trace runs past 200K tokens and keeps growing, pruning plus processed-token billing is the stronger design, and that is where Subconscious leads.

## What each one does

### Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### Crusoe

Crusoe started in 2018 turning wasted natural gas into power for computing and has since become a vertically integrated AI infrastructure company: it sources energy, builds data centers and rents GPUs through Crusoe Cloud. It designed and built the Abilene, Texas campus behind the OpenAI and Oracle Stargate project, planned at 1.2 GW, and in March 2026 announced an adjacent 900 MW campus for Microsoft. On September 17, 2026 it closed the first part of a $3.9B Series F at a $30.9B post-money valuation, and it reports over 6 GW of contracted capacity. Crusoe Cloud lists GB200 NVL72, B200 and AMD MI355X by quote, with H100 at $3.90 and H200 at $4.29 per GPU-hour on demand.

## Which is best, and when

### Choose Subconscious for

- Coding agents whose traces grow past 200K tokens
- Paying for processed tokens after compression
- Claude Code or Codex users wanting an open-model backend

### Choose Crusoe for

- Multi-turn chat that reuses long shared prefixes
- LoRA fine-tuning and serving in one account
- Raw GB200 or B200 clusters alongside inference

## At a glance

| Attribute | Subconscious | Crusoe |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | DeepSeek V4, GLM 5.3, Kimi K2.6, Nemotron 3 |
| Speed | 2x faster task completion | Up to 9.9x faster TTFT vs vLLM (vendor claim) |
| Price | 50–80% lower cost; billed on processed tokens | $0.05–$1.74 in, $0.20–$4.40 out per 1M |
| Customization | Marathon post-trained variants | Serverless LoRA fine-tuning |
| Deployment | Managed API, dedicated, on-prem | Serverless, self-serve and tailored dedicated, raw GPUs |
| Long context | 5M+ effective context | Varies by model; cluster-wide KV cache |

## FAQ

### What is the difference between Subconscious and Crusoe?

Both serve open models like GLM 5.3 and both target agents that resend long context. Crusoe reuses KV cache across a cluster; Subconscious prunes it and bills processed tokens.

### When should I choose Subconscious over Crusoe?

Coding agents whose traces grow past 200K tokens; Paying for processed tokens after compression; Claude Code or Codex users wanting an open-model backend.

### When should I choose Crusoe over Subconscious?

Multi-turn chat that reuses long shared prefixes; LoRA fine-tuning and serving in one account; Raw GB200 or B200 clusters alongside inference.

### Is Subconscious or Crusoe cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Crusoe: $0.05–$1.74 in, $0.20–$4.40 out per 1M. The cheaper choice depends on the model and workload.

### Which has more context, Subconscious or Crusoe?

Subconscious: 5M+ effective context. Crusoe: Varies by model; cluster-wide KV cache.

## Try Subconscious

Subconscious speaks the OpenAI and Anthropic API formats. Base URL: https://api.subconscious.dev/v1. Docs: https://docs.subconscious.dev. Get an API key: https://platform.subconscious.dev/signin. Agent guide: https://www.subconscious.dev/agents.md.

Full profiles: [Subconscious](https://www.subconscious.dev/providers/subconscious.md), [Crusoe](https://www.subconscious.dev/providers/crusoe.md).
