All posts
Announcements

Subconscious Raises $5.1 Million to Build the Inference Platform for Long-Running Agents

September 22, 2026

Jack O'Brien

Jack O'Brien

Co-Founder & CEO

Hongyin Luo

Hongyin Luo

Co-Founder & CTO

MIT researchers discovered a way to dynamically compress up to 80% of an AI agent’s context to make it much more efficient and accurate. Now, they’re turning that core technology into an opinionated inference platform to power agents that run can faster, for longer, at a lower cost with no changes to the underlying hardware, AI models, or apps that use them.

CAMBRIDGE, MA - September 22, 2026 - Subconscious is announcing $5.1 million in funds raised to launch an inference platform designed for long-running agents. Across a pre-seed and seed, MassVentures led the rounds, with participation from Foothill Ventures, Underscore VC, E14 Fund, Oakseed Ventures, and the Agent Fund among others. The inference platform is available to engineering teams today, with a cloud or on-prem deployment available for enterprises.

Language models are used in many ways: as a chatbot, as a classifier, as a document writer, or as an agent. Among all these uses, agents are extremely computationally intensive and require processing thousands of times as many tokens over long periods of time. Agents generate the most valuable work, but they’re most expensive way to use AI models. Already they’ve transformed software engineering, and they’re growing in usage among salespeople, lawyers, scientists, marketers, consultants, and all kinds of knowledge work. Based on recent trends, agents will consume virtually all inference in the very near future.

Subconscious built an inference platform to take advantage of the unique challenges in powering long-running agents. Born out of MIT research into inference optimizations, the system uses dynamic context compression and highly efficient caching for a stepwise gain in performance. For agents that consume beyond 200k tokens, Subconscious can complete tasks 2x faster, extend the effective context window of models to 5m+ tokens, improve **performance on key agentic benchmarks like coding and workflows by 1-10%, and decrease costs by up to 80%.

On the TriE benchmark, a benchmark built to measure system performance, Subconscious’s inference runtime completed tasks 2x faster than SGLang and supported 2.3x as many concurrent requests. Handling more concurrent requests means squeezing more work out of GPUs, effectively doubling or tripling the size of a cluster.

On DeepSWE, a benchmark built around long coding tasks, GLM 5.2 hosted on Subconscious solved 46% of problems at an average cost of $2.79. The same model served on standard inference infrastructure scored 44% and cost $3.92. Compression and caching from Subconscious improve both the model's efficiency and its accuracy.

Subconscious is live in production today and has been powering coding agents and agentic products in closed beta for months. In July, a 20-person engineering team switched from using Claude to using the GLM 5.2 model hosted on Subconscious to power their coding agents. In the two months since, they’ve cut their monthly AI spend from $40k to $6k, their engineers report faster token throughput, and they have yet to hit their rate limits. One engineer on the team ran an extremely long agent trace across 4,571 turns and 9,556 tool calls, and Subconscious recorded 449 million tokens where a conventional runtime would have billed 2.6 billion, an 82% reduction. Despite aggressive context compression, the engineering team reported no loss in model capability.

As AI costs have skyrocketed many companies have moved to cut costs, but engineers want more agents running for longer periods of time on harder problems. Subconscious allows companies to have it all.

"Open models finally got good enough this summer that their quality vs closed source models stopped being a compromise," says Jack O'Brien, Co-Founder and CEO of Subconscious. "Meanwhile every engineering team I talk to has put a ceiling on what it spends per developer. Teams need to spend less, but engineers are addicted and there’s no going back. We run open models in a way that's enhanced for coding agents, so teams can have it all."

"The Subconscious team is the best team on the planet to solve one of AI's toughest challenges," said Stacy Swider, VP of Investments at MassVentures. "We could not be more excited to be a part of their growth as they scale up to support thousands of companies and billions or even trillions of AI agents."

The Subconscious inference platform is live today, and teams can start powering coding agents like Claude Code, Codex, Pi, Copilot, and OpenCode with Subconscious in about 30 seconds with their easy CLI. Teams can also deploy the inference system on their own GPUs for maximum cost savings, fully air-gapped data protection, and ultimate control.

About Subconscious

Subconscious is an inference platform designed for long-horizon agents. Engineers choose Subconscious to power their agents to finish complex tasks faster, run for longer, improve their accuracy, and lower their costs substantially. The company is based in Cambridge, Massachusetts, and has raised $5.1 million led by MassVentures.

Subconscious aims to build the necessary infrastructure to power trillions of agents. Our mission is to allow everyone to do deep work and automate the rest without selling your soul.

Keep reading

All posts

Get started

Build your first agent.

Point your OpenAI or Anthropic SDK at our endpoint and try TIM in minutes.

Start building