Fast, low-cost inference for security workloads.
For security teams running long investigations and authorized testing. One API, 2x faster at 50–80% lower cost.

2x
Faster task completion
50–80%
Lower cost
Security work is token work.
An investigation pulls logs, correlates events across hosts, tests a theory and writes it up. Authorized testing runs just as long. Every step adds tokens, and the bill grows with the run.
Alert triage
Work every alert in the queue, not just the ones an analyst reaches before the shift ends.
Incident investigation
Correlate activity across hosts and accounts into one timeline with the evidence attached.
Detection engineering
Draft, test and tune detection rules against your own historical telemetry.
Authorized security testing
Scoped engagements for testing firms and internal red teams.
Step 40 depends on what step 12 found.
Findings chain together. Reset the context between steps and the agent loses the thread that connects four quiet alerts into one real incident.
5M+
Effective context
Runs keep their history past the model’s window, without lossy summaries.
Faster verdicts you can trust.
Security leaders weigh time to a verdict, confidence in the finding and where the data goes.
Faster verdicts
Long investigations return sooner, so analysts act while the answer still matters.
Findings that hold up
Less stale context gives the model a cleaner view, so late findings stay grounded in the evidence.
Nothing stored
The API keeps token counts for billing and nothing else. Logs and findings are never stored.
Where the tokens go.
Every investigation runs the same stages, and each one fills the context.
- 01
Collect
Pull logs, alerts and endpoint events. Raw telemetry is the bulk of the tokens.
- 02
Correlate
Link activity across hosts and accounts while every earlier finding stays in view.
- 03
Test
Check each theory against the evidence and drop the ones that fail.
- 04
Report
Write the verdict with every finding tied to its source.
Who builds this on Subconscious.
SOC tooling vendors
Adding agents that triage and investigate inside their platforms.
Security testing firms
Running authorized engagements for their clients.
Internal security teams
At large companies, putting agents on every alert in the queue.
Get started in minutes.
Every model is served in both the OpenAI and the Anthropic format. Point the agent you already have at Subconscious.
API format
Language
Endpoint
https://api.subconscious.dev/v1/chat/completions
Model
subconscious/glm-5.3-marathon
Works with
OpenAI SDK
Intact
Tool calls, streaming, structured output
from openai import OpenAI
client = OpenAI(
base_url="https://api.subconscious.dev/v1",
api_key="YOUR_API_KEY",
)
response = client.chat.completions.create(
model="subconscious/glm-5.3-marathon",
messages=[
{
"role": "user",
"content": "Refactor the billing module and update every affected test.",
}
],
)
print(response.choices[0].message.content)Questions
- We do not store your prompts or completions. The API keeps only token counts for billing. Teams with stricter rules can also deploy on their own infrastructure.
- The API serves GLM-5.3 Marathon and DeepSeek V4.1 Flash Marathon in both the OpenAI and Anthropic formats. Dedicated deployments can serve virtually any open model, including your own fine-tunes.
- Yes. Start on the Subconscious API, then move to a dedicated deployment in our cloud or on your own GPUs. The runtime is a drop-in replacement for vLLM or SGLang, and our engineers can install and tune it inside your cloud.
Need it on your own GPUs?
The same runtime runs as a dedicated deployment in our cloud or on your own hardware, as a drop-in replacement for vLLM or SGLang.