Inference for coding agents

Stop rationing your coding agent.

Get 1.8 billion tokens a month for $100. Use our API with your favorite coding agent in three lines of code.

zsh

# Install the Subconscious CLI

$npm install -g subconscious-cli

 

# Sign in, in your browser

$subc login

 

# Run Claude Code on Subconscious

$subc claude

  • 3.5x faster sustained performance on long traces
  • 5m+ context window
  • OpenAI and Anthropic compatible
  • Never trained on your code

Power your favorite harness with the Subconscious API

  • Claude Code
  • Codex
  • OpenCode
  • Cursor
  • Pi
  • Copilot

No new SDK, no rewritten prompts, no different tool format. The CLI points the harness you already use at Subconscious, and switches back just as fast if you hate it.

Coding plans

Enough tokens to stop thinking about tokens.

Get millions to billions of tokens per day. Run it down to zero and pick back up tomorrow, or add credits and keep going on usage based pricing.

Base

$100per month

Get moving fast on Subconscious

60M

tokens per day

1.8B

tokens per month

Pro

$500per month

For a small team or heavy users

300M

tokens per day

9B

tokens per month

Heavy

$2,000per month

For a team of heavy users

1.2B

tokens per day

36B

tokens per month

* Tokens are allocated per day and reset at midnight UTC. Daily allocation assumes usage in agent workloads, and a token mix of 98.2% cached tokens, 1.6% input tokens, and 0.2% output tokens, the weighting that characterizes long-horizon traces.

Why developers switch

Built for the way agents actually use tokens.

Powerful open models, enhanced for long context reasoning.

Headroom

No rate-limit roulette

Your daily allocation is yours. No weekly caps, no throttling three hours into a refactor, no waiting out someone else’s incident.

Unfiltered

No refusals on real work

An open model we serve ourselves, so legitimate work never trips a filter and quality never quietly degrades under load.

Privacy

Your code is never training data

Prompts and completions are never used to train anyone’s model. What your agents read stays yours.

Cost per agent run

Pay 6.3× less than Opus 5, and 2.9× less than standard GLM-5.3.

Standard inference vs. Subconscious with runtime context compression.

Less Tokens

More Tokens

Short

Med

Long

XL

Max

50K

300K

750K

1.5M

5M

Opus 5

$174.89

GLM-5.3

$80.69

GLM-5.3 Marathon

$27.56

Standard inferenceSubconscious
Retained context per step0500KContext window limitContext compression ceiling1750 stepsRetained context tokens

Effective context

750K tokens

Tokens billed

281.6M → 96.2M

Cost savings vs Opus

6.3× less

Cost savings vs standard GLM

2.9× less

Point your coding agent at us and see.

Three commands and you are running on your own allocation. Same harness, same workflow, and if it is not faster on your repo you are back where you started in a minute.

From developers

What people say after their first run.

It was faster than a lot of LLMs I've used, and it was easy to set up.

Cedric Prentice

Software Engineer @ Wayfair

Very cool, easy to use!

Bill Simmons

Co-Founder @ Orbit.me

Very interesting alternative to OpenAI

Dave Gogi

Founder @ Signal X

Pretty fast and fun to use

Nihir Kothari

Founder @ Sidekick Software

It worked very well and provided accurate descriptions of the images we passed in. The API costs were very cheap.

Sam Mayle

Researcher @ Mitsubishi Electric Research Lab

It was faster than a lot of LLMs I've used, and it was easy to set up.

Cedric Prentice

Software Engineer @ Wayfair

Very cool, easy to use!

Bill Simmons

Co-Founder @ Orbit.me

Very interesting alternative to OpenAI

Dave Gogi

Founder @ Signal X

Pretty fast and fun to use

Nihir Kothari

Founder @ Sidekick Software

It worked very well and provided accurate descriptions of the images we passed in. The API costs were very cheap.

Sam Mayle

Researcher @ Mitsubishi Electric Research Lab

It was faster than a lot of LLMs I've used, and it was easy to set up.

Cedric Prentice

Software Engineer @ Wayfair

Very cool, easy to use!

Bill Simmons

Co-Founder @ Orbit.me

Very interesting alternative to OpenAI

Dave Gogi

Founder @ Signal X

Pretty fast and fun to use

Nihir Kothari

Founder @ Sidekick Software

It worked very well and provided accurate descriptions of the images we passed in. The API costs were very cheap.

Sam Mayle

Researcher @ Mitsubishi Electric Research Lab

It was faster than a lot of LLMs I've used, and it was easy to set up.

Cedric Prentice

Software Engineer @ Wayfair

Very cool, easy to use!

Bill Simmons

Co-Founder @ Orbit.me

Very interesting alternative to OpenAI

Dave Gogi

Founder @ Signal X

Pretty fast and fun to use

Nihir Kothari

Founder @ Sidekick Software

It worked very well and provided accurate descriptions of the images we passed in. The API costs were very cheap.

Sam Mayle

Researcher @ Mitsubishi Electric Research Lab

Made the whole process very easy!

Smruthi Ramesh

Lead Data Scientist @ Schneider Electric

It was really fast!!

Inder Singh

UDE

Easy to use, great UI

Yassine Fatimi

Founder @ ClauseGuard

Pretty easy to use. No brainer. Easy drop in for OpenAI.

Hansen Liang

Founder @ stealth

Made the whole process very easy!

Smruthi Ramesh

Lead Data Scientist @ Schneider Electric

It was really fast!!

Inder Singh

UDE

Easy to use, great UI

Yassine Fatimi

Founder @ ClauseGuard

Pretty easy to use. No brainer. Easy drop in for OpenAI.

Hansen Liang

Founder @ stealth

Made the whole process very easy!

Smruthi Ramesh

Lead Data Scientist @ Schneider Electric

It was really fast!!

Inder Singh

UDE

Easy to use, great UI

Yassine Fatimi

Founder @ ClauseGuard

Pretty easy to use. No brainer. Easy drop in for OpenAI.

Hansen Liang

Founder @ stealth

Made the whole process very easy!

Smruthi Ramesh

Lead Data Scientist @ Schneider Electric

It was really fast!!

Inder Singh

UDE

Easy to use, great UI

Yassine Fatimi

Founder @ ClauseGuard

Pretty easy to use. No brainer. Easy drop in for OpenAI.

Hansen Liang

Founder @ stealth

On-prem and dedicated

Want Subconscious inside your cloud?

For full control, run our inference system on your own GPUs. Use any model and harness, and increase the capacity of your GPU cluster.

Talk to us about on-prem
Models
GLM, Nemotron, or any open weight model
Compatibility
Any coding harness
Capacity
Billions of tokens per day per node
Pricing
Per-GPU under management