Fast, low-cost inference for hedge fund research.
For research agents that read, model and backtest all day. Long runs finish 2x faster at 50–80% lower cost, and we never store your prompts.

2x
Faster task completion
50–80%
Lower cost
Funds run on tokens now.
Agents read filings and transcripts, pull numbers into models, backtest, revise and maintain the systems behind it all. A fund can run thousands of these sessions a day, and every one runs long.
Filing and transcript research
Pull the numbers from hundreds of filings and calls into one model.
Backtest and revise loops
Test a strategy, read why it failed and try the next version.
Research infrastructure
Maintain the data pipelines and signal code your researchers depend on.
Thesis monitoring
Watch new filings and news against the positions you already hold.
The value is remembering what already failed.
A backtest-and-revise loop is many steps by nature. The agent earns its keep by remembering what it tried and why it didn’t work.
5M+
Effective context
Runs keep their history past the model’s window, without lossy summaries.
Speed is edge. So is secrecy.
A fund wins by testing more ideas sooner than everyone else, without anyone else seeing them.
Results before the market moves
Long research runs return sooner, so ideas get tested while they still matter.
More bets per dollar
Most research runs lead nowhere. Lower cost per run means more of them in parallel.
Nothing stored
The API keeps token counts for billing and nothing else. Your research is never stored.
Your research budget covers far more work.
50–80% lower cost means the same spend buys two to five times the research. Run more theses in parallel, longer backtests or a second pass to check the first.
2–5×
more research per dollar
What we never store.
The API keeps token counts for billing. None of this is ever stored.
- Prompts and research notes
- Positions and portfolio data
- Strategy code and signals
- Backtest results
- Fine-tuned weights
- Logs and traces
Who builds this on Subconscious.
Hedge funds
Running research agents at volume, next to their own data.
Quant and prop shops
Testing many strategies cheaply, in parallel.
Fund engineering teams
Building the internal agent platform their researchers use.
Get started in minutes.
Every model is served in both the OpenAI and the Anthropic format. Point the agent you already have at Subconscious.
API format
Language
Endpoint
https://api.subconscious.dev/v1/chat/completions
Model
subconscious/glm-5.3-marathon
Works with
OpenAI SDK
Intact
Tool calls, streaming, structured output
from openai import OpenAI
client = OpenAI(
base_url="https://api.subconscious.dev/v1",
api_key="YOUR_API_KEY",
)
response = client.chat.completions.create(
model="subconscious/glm-5.3-marathon",
messages=[
{
"role": "user",
"content": "Refactor the billing module and update every affected test.",
}
],
)
print(response.choices[0].message.content)Questions
- No. The API keeps only token counts for billing, never prompts or completions. Funds that want more can also deploy on their own GPUs.
- Sign up, mint a key and change the base URL. The API speaks both the OpenAI and Anthropic formats, so existing research agents run as they are.
- Yes. Start on the Subconscious API, then move to a dedicated deployment in our cloud or on your own GPUs. The runtime is a drop-in replacement for vLLM or SGLang, and our engineers can install and tune it inside your cloud.
Need it on your own GPUs?
The same runtime runs as a dedicated deployment in our cloud or on your own hardware, as a drop-in replacement for vLLM or SGLang.