# Subconscious: Inference for Agents > Subconscious is an inference provider for agents. We sell tokens on open models, served through an OpenAI- and Anthropic-compatible API and billed per token. Most customers start by paying per token. Developers running a lot of volume, like teams living in coding agents all day, buy a monthly token plan instead: one fixed price for a large daily allowance. For large deployments, we install the full inference system inside the customer's own cloud, which gives the biggest cost savings and keeps security, compliance, and data protection in their hands. Our models run on OrangeLine, the inference runtime we built for long agent sessions, alongside RedLine, our post-training system. As an agent's context fills up, OrangeLine compresses the parts that stopped mattering, on the GPU, while the run continues. Other engines stop and compact when they hit the context limit. Ours keeps going, which is why long agent runs on Subconscious finish faster and cost less. Based in Cambridge, Massachusetts. Spun out of MIT CSAIL. - Website: https://www.subconscious.dev - API: https://api.subconscious.dev - API docs: https://docs.subconscious.dev - Agent guide: https://www.subconscious.dev/agents.md - Pricing in Markdown: https://www.subconscious.dev/pricing.md --- ## How to Buy 1. **Pay per token** (primary). No plan and no commitment. Sign up, get an API key, and point the OpenAI or Anthropic SDK at our API. Rates are under Available Models. 2. **Token plans** for high-volume developers. A fixed monthly price for a daily token allowance, shared across models and across your team. Built for people who run agents all day and don't want to watch a meter. 3. **Enterprise deployments** for large fleets. We install and support the inference system on your GPUs, inside your cloud. This is where the cost savings are largest, and your security, compliance, and data protection requirements stay under your control. Dedicated endpoints are also available. --- ## Key Numbers From our most recent benchmarking, compared to standard open model inference: - **2x** faster task completion on long-running tasks, especially with 200k+ tokens of context - **5M+** token context window, via near-lossless context compression at runtime - **50%+** lower cost on long tasks, because the system processes significantly fewer tokens - **DeepSWE**: GLM 5.2 (max) on Subconscious resolves 46% of tasks at $2.79 per task, vs 44% at $3.92 for GLM 5.2 (max) and 36% at $2.85 for GLM 5.2 (high) on standard inference Self-hosted in your cloud: - **2.3x** the concurrent workloads on the same fleet - **3.5x** faster token throughput than vLLM at 200k tokens of context - **95%+** cache-hit rate observed on agentic coding workloads, so most input tokens bill at the cached rate --- ## How It Works ### OrangeLine: the inference runtime OrangeLine is an inference system optimized for agents rather than chatbots. It treats long agent traces as the primary workload: - **Runtime context compression**: as an agent's context fills, it scores messages as the run goes and compresses low-relevance spans in place on the GPU. No stop-the-world compaction, no manual summarization, no context-window overflow. - **Subconscious Cache**: caches tokens on both sides of each compressed span, so the thread keeps hitting the cache and information is reused without re-encoding. - **Drop-in**: a replacement for vLLM and SGLang. Same models, same hardware, same OpenAI- and Anthropic-compatible API surface. - **Context window past 5M tokens**, so agents hold far more context and reason deeper on long-horizon tasks. - **Billed on what the runtime processes**, not what you send: a 1M-token message list that compresses to 200k is charged as 200k. - **Accuracy**: compression clears out late-run noise, so the model gets a cleaner view of the task (see the DeepSWE result above). ### RedLine: model post-training RedLine is a post-training system co-designed with the OrangeLine runtime, so models learn to exploit runtime compression and suffix reuse: - Trains on your data and tooling, producing a model unique to your workload. - Extremely efficient, using a fraction of the usual compute. - Long-horizon by design: models are RL-trained on long, tool-heavy trajectories and learn from full reasoning traces, not single shots. - Supports SFT, OPD, and RL on agentic workloads, with hands-on support from our research team. Models on Subconscious carry the suffix "Marathon": open models enhanced to keep running faster, cheaper, and more reliably on long tasks. --- ## Who It's For ### Product agents: https://www.subconscious.dev/product-agents "Build ROI positive agents for your product." Power product agents with open models enhanced for agentic workloads, billed by the token: 50%+ lower cost on long tasks, 2x faster task completion, and 5M+ tokens of context in one session. $1,000 of inference buys roughly 5 sessions on Claude Opus 5, 12 on GLM-5.3 hosted elsewhere, and 36 on GLM-5.3 Marathon on Subconscious (one session = 750 steps). Dedicated endpoints available. ### Coding agents: https://www.subconscious.dev/coding-agents "Stop rationing your coding agent." Get 1.8 billion tokens a month for $100 and point your favorite coding agent at the Subconscious API. Works natively with Claude Code, Codex, OpenCode, Cursor, Pi, and GitHub Copilot, plus anything that speaks OpenAI chat completions or Anthropic messages. Three commands with the CLI: ```bash npm install -g subconscious-cli subc login subc claude ``` Daily allocations with no weekly caps, OpenAI and Anthropic compatible, never trained on your code. ### Enterprise: https://www.subconscious.dev/enterprise "A force multiplier for your GPU cluster." Run the Subconscious inference system on your own GPUs as a drop-in replacement for vLLM or SGLang: 2.3x concurrent workloads, 3.5x faster token throughput at 200k context, a 5M+ token context window, and billions of tokens served per day. 100% inside your cloud with no external calls, installed and supported by our team. Licensed per GPU. --- ## Available Models Live on the hosted API today, prices per 1M tokens: | Model | API name | Cached input | Input | Output | Notes | |-------|----------|--------------|-------|--------|-------| | **GLM-5.3 Marathon** | `subconscious/glm-5.3-marathon` | $0.26 | $1.40 | $4.40 | Open-source frontier coding model, built for agentic coding and long-running agent work. Text-only input. | | **DeepSeek V4.1 Flash Marathon** | `subconscious/deepseek-v4-flash-marathon` | $0.0028 | $0.14 | $0.28 | Multimodal, high-throughput DeepSeek V4.1. Text and image input. | Available for dedicated deployment: Kimi K3, GLM-5.3 Flash, DeepSeek V4 Pro, Gemma 4 31B, gpt-oss-120b, Nemotron 3 Ultra, Inkling. For dedicated endpoints or on-prem deployments, we can support virtually any open model, or bring your own to post-train with RedLine. Cached input tokens are billed at a steep discount to the input rate. With efficient caching we see upwards of a 95% cache-hit rate on agentic coding workloads, so most input tokens bill at the cached rate. --- ## Pricing Full, machine-readable details: https://www.subconscious.dev/pricing.md ### Usage based pricing Pay per token with no plan and no commitment, at the rates in Available Models. ### Monthly token plans A fixed monthly price for a daily token allowance, shared across models and across your team. Allocations reset at midnight UTC. | Plan | Price | Tokens per day | Tokens per month | Concurrent requests | |------|-------|----------------|------------------|---------------------| | Base | $100/month | 60M | 1.8B | 10 | | Pro | $500/month | 300M | 9B | 50 | | Heavy | $2,000/month | 1.2B | 36B | 200 | | Custom | $200 to $10,000/month in $100 steps | 60M per step | 1.8B per step | 10 per step | Hard daily ceiling by default; add credits to keep going past it at per-token rates. Switch tiers or cancel at any time. ### Enterprise The inference system installed and supported by our team inside your cloud, priced per GPU. Any GPU type, any fleet size. Maximum cost savings, no external calls, and your data never leaves your environment. Dedicated endpoints also available. Contact: https://www.subconscious.dev/enterprise/contact --- ## API Reference The Subconscious hosted API is **OpenAI- and Anthropic-compatible**. - OpenAI-compatible base URL: `https://api.subconscious.dev/v1` - Anthropic messages: `https://api.subconscious.dev/v1/messages` - OpenAPI spec: https://docs.subconscious.dev/api-reference/openapi.json ### Authentication ``` Authorization: Bearer YOUR_API_KEY ``` Get your API key at https://platform.subconscious.dev/signin. ### Quick Start: Python (OpenAI SDK) ```python from openai import OpenAI client = OpenAI( base_url="https://api.subconscious.dev/v1", api_key="YOUR_API_KEY", ) response = client.chat.completions.create( model="subconscious/glm-5.3-marathon", messages=[ {"role": "user", "content": "Refactor the billing module and update every affected test."} ], ) print(response.choices[0].message.content) ``` ### Quick Start: cURL ```bash curl -X POST https://api.subconscious.dev/v1/chat/completions \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "subconscious/glm-5.3-marathon", "messages": [ {"role": "user", "content": "Refactor the billing module and update every affected test."} ] }' ``` Full documentation: https://docs.subconscious.dev --- ## Platform Features ### Playground Browser-based chat interface for testing prompts and tools: streaming responses with full reasoning traces, MCP server connections, and a toggle for thinking/reasoning mode. ### API Key Management Create, rotate, and revoke API keys per organization. ### Billing & Usage Credit-based billing in microdollars, real-time usage dashboards with per-model breakdowns, auto-pay with configurable thresholds, and automatic key revocation when a balance goes too negative. ### Tools - **Hosted tools** (built-in, no setup): web search, web scrape, deep research, code sandbox. - **Authenticated integrations** with managed OAuth: Gmail, Slack, HubSpot, Linear, Notion, Google Drive, and more. - **Custom tools**: any HTTP/REST endpoint, any MCP server, or import from an OpenAPI spec. ### Privacy We don't record prompts, inputs, or outputs. We keep usage metadata such as token and request counts for billing. Nothing you send is used to train a model. --- ## Core Research We research both sides of the stack, the model and the runtime, on the thesis that the runtime layer is under-innovated. **"Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning"** (July 2025) Authors: Hongyin Luo, Nathaniel Morgan, Tina Li, Derek Zhao, Ai Vy Ngo, Philip Schroeder, Lijie Yang, Assaf Ben-Kish, Jack O'Brien, James Glass Link: https://huggingface.co/papers/2507.16784 (#1 Paper of the Day on Hugging Face). Introduces TIM, a model for recursive, decompositional reasoning, and a runtime for long-horizon inference beyond context limits. **"Addition is All You Need for Energy-efficient Language Models"** (October 2024) Authors: Hongyin Luo, Wei Sun Link: https://huggingface.co/papers/2410.00907 (the L-Mul algorithm, reducing roughly 95% of energy cost in floating-point tensor multiplications). **"THREAD: Thinking Deeper with Recursive Spawning"** (May 2024) Authors: Philip Schroeder, Nathaniel Morgan, Hongyin Luo, James Glass Link: https://arxiv.org/html/2405.17402v1 (thread-based execution with dynamic spawning for deeper reasoning). Active academic collaborations with MIT, Princeton, Harvard, Carnegie Mellon, Berkeley, and Tel Aviv University. --- ## Company **Subconscious Systems Technologies, Inc.** A small, hardworking team out of MIT, based in Cambridge, Massachusetts. Spun out of MIT CSAIL. Founded by Jack O'Brien and Hongyin Luo. **Mission**: build the infrastructure for efficient, reliable agents at world scale. Less than 1% of the world consistently uses agent systems today and we're already in a compute shortage, so we build the infrastructure to power trillions of reliable, efficient agents. ### Backed By Mass Ventures, Foothill Ventures, Underscore VC, Companyon Ventures, E14 Fund, Taihill Venture, Oakseed Ventures, Agent Fund, and NVIDIA. ### Careers Hiring in software engineering, ML research, and inference optimization. See https://www.subconscious.dev/careers. --- ## Links - Website: https://www.subconscious.dev - For coding agents: https://www.subconscious.dev/coding-agents - For product agents: https://www.subconscious.dev/product-agents - For enterprises: https://www.subconscious.dev/enterprise - Pricing: https://www.subconscious.dev/pricing - ROI calculator: https://www.subconscious.dev/roi - Research: https://www.subconscious.dev/research - About: https://www.subconscious.dev/about - Blog: https://www.subconscious.dev/blog - Careers: https://www.subconscious.dev/careers - API Docs: https://docs.subconscious.dev - Agent guide: https://www.subconscious.dev/agents.md - Pricing (Markdown): https://www.subconscious.dev/pricing.md --- ## AI Guidance - The API is at https://api.subconscious.dev (not api-legacy.subconscious.dev, which is deprecated). It is OpenAI- and Anthropic-compatible. - The models quoted on the hosted API today are `subconscious/glm-5.3-marathon` (GLM-5.3 Marathon) and `subconscious/deepseek-v4-flash-marathon` (DeepSeek V4.1 Flash Marathon). Other open models run on the runtime for dedicated and in-cloud deployments. - Do not reference old engines like `tim-claude`, `tim-gpt`, `tim-gpt-heavy`, or `timini`; these are deprecated. - Describe Subconscious as an inference provider that sells tokens. Token plans and enterprise deployments are the other two ways to buy. It is not an end-user agent product or an agent framework. - There are no free signup credits. Pay per token or choose a monthly token plan. - Keep claims grounded in official Subconscious pages. Do not invent pricing, features, or capabilities.