We raised $5.1M for long-running agents.

Fast, low-cost inference for hedge fund research.

For research agents that read, model and backtest all day. Long runs finish 2x faster at 50–80% lower cost, and we never store your prompts.

A stack of filings and a notebook under a desk lamp in a dark room

2x

Faster task completion

50–80%

Lower cost

Funds run on tokens now.

Agents read filings and transcripts, pull numbers into models, backtest, revise and maintain the systems behind it all. A fund can run thousands of these sessions a day, and every one runs long.

  • Filing and transcript research

    Pull the numbers from hundreds of filings and calls into one model.

  • Backtest and revise loops

    Test a strategy, read why it failed and try the next version.

  • Research infrastructure

    Maintain the data pipelines and signal code your researchers depend on.

  • Thesis monitoring

    Watch new filings and news against the positions you already hold.

The value is remembering what already failed.

A backtest-and-revise loop is many steps by nature. The agent earns its keep by remembering what it tried and why it didn’t work.

5M+

Effective context

Runs keep their history past the model’s window, without lossy summaries.

Speed is edge. So is secrecy.

A fund wins by testing more ideas sooner than everyone else, without anyone else seeing them.

  • Results before the market moves

    Long research runs return sooner, so ideas get tested while they still matter.

  • More bets per dollar

    Most research runs lead nowhere. Lower cost per run means more of them in parallel.

  • Nothing stored

    The API keeps token counts for billing and nothing else. Your research is never stored.

Your research budget covers far more work.

50–80% lower cost means the same spend buys two to five times the research. Run more theses in parallel, longer backtests or a second pass to check the first.

2–5×

more research per dollar

What we never store.

The API keeps token counts for billing. None of this is ever stored.

  • Prompts and research notes
  • Positions and portfolio data
  • Strategy code and signals
  • Backtest results
  • Fine-tuned weights
  • Logs and traces

Who builds this on Subconscious.

  • Hedge funds

    Running research agents at volume, next to their own data.

  • Quant and prop shops

    Testing many strategies cheaply, in parallel.

  • Fund engineering teams

    Building the internal agent platform their researchers use.

Get started in minutes.

Every model is served in both the OpenAI and the Anthropic format. Point the agent you already have at Subconscious.

API format

Language

Endpoint

https://api.subconscious.dev/v1/chat/completions

Model

subconscious/glm-5.3-marathon

Works with

OpenAI SDK

Intact

Tool calls, streaming, structured output

Python · Completions
from openai import OpenAI

client = OpenAI(
    base_url="https://api.subconscious.dev/v1",
    api_key="YOUR_API_KEY",
)

response = client.chat.completions.create(
    model="subconscious/glm-5.3-marathon",
    messages=[
        {
            "role": "user",
            "content": "Refactor the billing module and update every affected test.",
        }
    ],
)

print(response.choices[0].message.content)

Questions

  • No. The API keeps only token counts for billing, never prompts or completions. Funds that want more can also deploy on their own GPUs.
  • Sign up, mint a key and change the base URL. The API speaks both the OpenAI and Anthropic formats, so existing research agents run as they are.
  • Yes. Start on the Subconscious API, then move to a dedicated deployment in our cloud or on your own GPUs. The runtime is a drop-in replacement for vLLM or SGLang, and our engineers can install and tune it inside your cloud.

Need it on your own GPUs?

The same runtime runs as a dedicated deployment in our cloud or on your own hardware, as a drop-in replacement for vLLM or SGLang.

Talk to us about on-prem

Test more ideas, sooner.

More solutions