We raised $5.1M for long-running agents.

Run your team’s coding agents 2x faster at half the cost.

For engineering teams running agents at every developer’s desk, in CI and in the cloud. One API, compatible with the agents you already use.

Cost of one agent run
$707Standard$200Subconscious05M tokens

72% less at 5M tokens. GLM-5.3 at list rates.

The numbers, on long coding runs.

Measured against the same open models on standard inference, at 200k+ tokens of context.

Faster task completion

2x

Measured at 200k+ tokens of context, same model and hardware.
Lower cost

50%+

The runtime compresses spans before they are billed, so you pay for what it processes.
Accuracy gain

1-10%

On agentic benchmarks. Less stale context gives the model a cleaner view of the task.
Effective context

5M+

Runs keep their history past the model’s window, without lossy summaries.

More accurate, for less, on DeepSWE.

On the DeepSWE benchmark for long-horizon agentic coding, GLM-5.2 scores 46% on Subconscious against 44% on standard inference, at a lower cost per task.

DeepSWE (Long-Horizon Agentic Coding Benchmark)

36%40%44%$3.00$3.50$4.00Cost per task →Accuracy →GLM 5.2 (high)$2.85 · 36%GLM 5.2 (max)$3.92 · 44%GLM 5.2 (max) on Subconscious$2.79 · 46%

Coding is the longest-running agent work there is.

A coding agent reads the repo, plans, edits, runs the tests and tries again. Sessions pass 200k tokens fast, and background agents run for hours with nobody watching. Every step costs time and money.

  • Coding agents at every desk

    Claude Code, Codex, OpenCode and Cursor, pointed at Subconscious with one config change.

  • Background agents

    Bug fixes, migrations and test backfills that run while the team sleeps.

  • Code review at scale

    Review every pull request with the whole diff and the codebase around it in view.

  • CI repair

    Agents that read the failing build, fix it and push.

The plan from step 2 has to survive step 2,000.

Exploration fills the context with search results and tool output. If the agent loses its plan in that noise, it repeats work or ships the wrong fix.

Throughput is the product.

For a team running agents all day, more tokens through the system means more work shipped. We move the two numbers that limit it.

  • Results sooner

    Generation stays fast as the context fills, so long tasks return while they still matter.

  • A budget that stretches

    The runtime compresses stale context before it is billed, so exploring stops being the expensive part.

  • No rewrite

    OpenAI and Anthropic compatible. Change the base URL and your agents run as they are.

Works with the coding agents your team already uses.

Point any OpenAI- or Anthropic-compatible agent at Subconscious. Setup is a base URL, a key and a model string.

Coding Plans for your teamA daily token allocation for every developer, at a fixed monthly price.See coding plans

Built for background agents.

A cloud agent repeats four steps until the task holds. Marathon Mode makes each one faster and cheaper.

  1. 01

    Explore

    Search results and tool output pile up fast. The runtime compresses what stopped mattering before it is billed.

  2. 02

    Plan

    The plan stays in context for the whole task, so the agent never loses the thread.

  3. 03

    Execute

    Generation stays fast as context grows, so long tasks keep moving.

  4. 04

    Verify

    A cleaner context gives the model a sharper view when it checks its own work.

Your budget covers far more agent work.

50–80% lower cost means the same spend buys two to five times the work. Put it into more runs, longer runs or extra verification passes.

2–5×

more agent work per dollar

Who builds this on Subconscious.

  • Engineering teams

    Running coding agents on every developer’s desk and in CI.

  • Agent platforms

    Running background agents for their users and paying for every token.

  • Dev-tool companies

    Shipping cloud coding agents that work while engineers sleep.

Get started in minutes.

Every model is served in both the OpenAI and the Anthropic format. Point the agent you already have at Subconscious.

API format

Language

Endpoint

https://api.subconscious.dev/v1/chat/completions

Model

subconscious/glm-5.3-marathon

Works with

OpenAI SDK

Intact

Tool calls, streaming, structured output

Python · Completions
from openai import OpenAI

client = OpenAI(
    base_url="https://api.subconscious.dev/v1",
    api_key="YOUR_API_KEY",
)

response = client.chat.completions.create(
    model="subconscious/glm-5.3-marathon",
    messages=[
        {
            "role": "user",
            "content": "Refactor the billing module and update every affected test.",
        }
    ],
)

print(response.choices[0].message.content)

Questions

  • Yes. Point any agent or harness that speaks the OpenAI or Anthropic API at Subconscious by changing the base URL, key and model string. Setup guides cover Claude Code, Codex, OpenCode, Cursor and more.
  • Yes. Coding Plans give each developer on your team a daily token allocation for any coding agent or harness, at a fixed monthly price.
  • Yes. Start on the shared API, then move to a dedicated deployment with its own GPUs and no rate limits as volume grows.
  • The API serves GLM-5.3 Marathon and DeepSeek V4.1 Flash Marathon in both the OpenAI and Anthropic formats. Dedicated deployments can serve virtually any open model, including your own fine-tunes.

Need it on your own GPUs?

The same runtime runs as a dedicated deployment in our cloud or on your own hardware, as a drop-in replacement for vLLM or SGLang.

Talk to us about dedicated

Ship more with every agent run.

More solutions