# Subconscious

> Excellent speed, cost, and accuracy on tasks that need 200k+ tokens.

Canonical: https://www.subconscious.dev/providers/subconscious · By The Subconscious Team · Updated September 30, 2026

- Founded: 2025
- Example models: GLM 5.3, DeepSeek V4.1 Flash
- Website: https://www.subconscious.dev

## Overview

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

That design changes the bill. Subconscious charges for tokens its system actually processes after compression, not tokens sent, so a request that sends 1M tokens might bill for 200K. The managed API serves GLM 5.3 and DeepSeek V4.1 Flash, and dedicated or on-prem deployments can run nearly any open model. It speaks the OpenAI and Anthropic SDK formats and plugs straight into Claude Code, Codex, Cursor, GitHub Copilot and OpenCode. Subconscious records no prompts or inputs, only usage data.

## Upsides

- Speed, cost and accuracy that improve as context grows past 200K tokens, where most hosts get slower and pricier.
- Billing on processed tokens rewards the long, cache-heavy traces coding agents produce.

## Core use cases

- Coding agents working over 200k tokens.
- Research, review and multi-step enterprise agents.
- Agentic user-facing products across domains.

## Downsides

- A focused catalog of a few open models on the managed API.
- Short, single-turn requests see little of the advantage, since the gains come from long traces.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open weights |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash |
| Speed | 2x faster task completion |
| Price | 50–80% lower cost; billed on processed tokens |
| Customization | Marathon post-trained variants |
| Deployment | Managed API, dedicated, on-prem |
| Long context | 5M+ effective context |

## FAQ

### What is Subconscious?

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### What is Subconscious best for?

Coding agents working over 200k tokens; Research, review and multi-step enterprise agents; Agentic user-facing products across domains.

### How much does Subconscious cost?

Subconscious pricing at a glance: 50–80% lower cost; billed on processed tokens. Rates change often, so check Subconscious's pricing page before committing.

### How much context does Subconscious support?

Subconscious's long-context support: 5M+ effective context.

### What are the downsides of Subconscious?

A focused catalog of a few open models on the managed API; Short, single-turn requests see little of the advantage, since the gains come from long traces.

### What are the best alternatives to Subconscious?

Common alternatives include OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock, Together AI. Each has a head-to-head comparison with Subconscious on this site.

## Comparisons

- [Subconscious vs OpenAI](https://www.subconscious.dev/compare/subconscious-vs-openai.md)
- [Subconscious vs Anthropic](https://www.subconscious.dev/compare/subconscious-vs-anthropic.md)
- [Subconscious vs Google Vertex AI](https://www.subconscious.dev/compare/subconscious-vs-google-vertex.md)
- [Subconscious vs Amazon Bedrock](https://www.subconscious.dev/compare/subconscious-vs-aws-bedrock.md)
- [Subconscious vs Together AI](https://www.subconscious.dev/compare/subconscious-vs-together-ai.md)
- [Subconscious vs Fireworks AI](https://www.subconscious.dev/compare/subconscious-vs-fireworks.md)
- [Subconscious vs Baseten](https://www.subconscious.dev/compare/subconscious-vs-baseten.md)
- [Subconscious vs Groq](https://www.subconscious.dev/compare/subconscious-vs-groq.md)
- [Subconscious vs Cerebras](https://www.subconscious.dev/compare/subconscious-vs-cerebras.md)
- [Subconscious vs DeepInfra](https://www.subconscious.dev/compare/subconscious-vs-deepinfra.md)
- [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md)
- [Subconscious vs xAI](https://www.subconscious.dev/compare/subconscious-vs-xai.md)
- [Subconscious vs DeepSeek](https://www.subconscious.dev/compare/subconscious-vs-deepseek.md)
- [Subconscious vs Moonshot AI](https://www.subconscious.dev/compare/subconscious-vs-moonshot-ai.md)
- [Subconscious vs Z.ai](https://www.subconscious.dev/compare/subconscious-vs-z-ai.md)
- [Subconscious vs Alibaba Cloud](https://www.subconscious.dev/compare/subconscious-vs-alibaba-cloud.md)
- [Subconscious vs Meta](https://www.subconscious.dev/compare/subconscious-vs-meta.md)
- [Subconscious vs SambaNova](https://www.subconscious.dev/compare/subconscious-vs-sambanova.md)
- [Subconscious vs Nebius](https://www.subconscious.dev/compare/subconscious-vs-nebius.md)
- [Subconscious vs fal](https://www.subconscious.dev/compare/subconscious-vs-fal.md)
- [Subconscious vs Novita AI](https://www.subconscious.dev/compare/subconscious-vs-novita-ai.md)
- [Subconscious vs Parasail](https://www.subconscious.dev/compare/subconscious-vs-parasail.md)
- [Subconscious vs Inference.net](https://www.subconscious.dev/compare/subconscious-vs-inference-net.md)
- [Subconscious vs GMI Cloud](https://www.subconscious.dev/compare/subconscious-vs-gmi-cloud.md)
- [Subconscious vs Sail Research](https://www.subconscious.dev/compare/subconscious-vs-sail-research.md)
- [Subconscious vs Morph](https://www.subconscious.dev/compare/subconscious-vs-morph.md)
- [Subconscious vs Relace](https://www.subconscious.dev/compare/subconscious-vs-relace.md)
- [Subconscious vs TypeSafe AI](https://www.subconscious.dev/compare/subconscious-vs-typesafe-ai.md)
- [Subconscious vs StepFun](https://www.subconscious.dev/compare/subconscious-vs-stepfun.md)
- [Subconscious vs Runware](https://www.subconscious.dev/compare/subconscious-vs-runware.md)
- [Subconscious vs StreamLake](https://www.subconscious.dev/compare/subconscious-vs-streamlake.md)
- [Subconscious vs Wafer](https://www.subconscious.dev/compare/subconscious-vs-wafer.md)
- [Subconscious vs RunInfra](https://www.subconscious.dev/compare/subconscious-vs-runinfra.md)
- [Subconscious vs Particle.AI](https://www.subconscious.dev/compare/subconscious-vs-particle-ai.md)

## Sources

- [Subconscious](https://subconscious.dev)

Pricing and model lineups change often; figures are a snapshot.
