Pricing

Usage based pricing or monthly plans to power your agents efficiently.

Monthly token plans

Tokens nearly too cheap to meter

Pick a daily allocation and share it across as many people as you want. No per-seat pricing, no per-request metering. Use it for coding or your product.

Base

Get moving fast on Subconscious

$100/ month

60Mtokens per day

Tokens per month1.8B
Concurrent requests10
  • Every model on the API
  • Drop-in for Claude Code, Codex, OpenCode, or any OpenAI-compatible harness
Get Base

Cancel any time

Most popular

Pro

For a small team or heavy users

$500/ month

300Mtokens per day

Tokens per month9B
Concurrent requests50
  • Everything in Base, with 5x the daily allocation
  • Enough concurrency to fan out agents across a large codebase
Get Pro

Cancel any time

Heavy

For a team of heavy users

$2,000/ month

1.2Btokens per day

Tokens per month36B
Concurrent requests200
  • Everything in Pro, with 4x the daily allocation
  • Per-member API keys and a shared usage dashboard
  • Direct line to our engineering team
Get Heavy

Cancel any time

Build your own

Custom

Dial in the exact allocation your team burns through

$5,000/ mo

$200 to $10,000, in $100 steps

Tokens a day3B
Tokens per month90B
Concurrent requests500
  • Any allocation between the bounds, in $100 steps
  • Per-member API keys and a shared usage dashboard
  • Direct line to our engineering team
Get Custom at $5,000

Cancel any time

Every plan includes

  • Use it solo, or invite your team and share the daily allocation between you
  • Hard daily ceiling by default, so a runaway agent cannot surprise you
  • Add credits to keep going past the ceiling at the per-token rates below
  • Switch tiers or cancel at any time from the dashboard

Ceilings reset daily, so the monthly figure assumes a 30 day month. Tokens per month assume a token mix of 98.2% cached, 1.6% input, and 0.2% output, characteristic of long-horizon tasks. Need more than 6B tokens a day, or a dedicated endpoint? Talk to us about a deployment in your cloud below.

Usage based pricing

Or pay only for the tokens you use

Pay per token with no plan and no commitment. Working past your token plan's ceiling? Tack on additional tokens at these standard rates.

ModelCached tokensInput tokensOutput tokens
GLM-5.3 Marathon

Open-source frontier coding model

subconscious/glm-5.3-marathon

$0.26$1.40$4.40
DeepSeek V4 Flash Marathon

Multimodal & high-throughput DeepSeek V4

subconscious/deepseek-v4-flash-marathon

$0.0028$0.14$0.28
Kimi K3
Available for dedicated deploymentContact us
GLM-5.3 Flash
Available for dedicated deploymentContact us
DeepSeek V4 Pro
Available for dedicated deploymentContact us
Gemma 4 31B
Available for dedicated deploymentContact us
gpt-oss-120b
Available for dedicated deploymentContact us
Nemotron 3 Ultra
Available for dedicated deploymentContact us
Inkling
Available for dedicated deploymentContact us

Prices in USD per 1M tokens. With efficient caching, we see upwards of a 95% cache hit rate on agentic coding workloads, so most input tokens bill at the cached rate.

What 9B tokens per month costs

Power your agents with more tokens, for a lot less

Subconscious Pro plan

$500

Flat monthly fee, shared across your team

Subconscious, pay per token

$2,579

Billed per request

GLM-5.3 on other platforms

$3,610

Identical list price, but without context compression

Claude Opus 5

$7,825

Frontier list price

Assumes agentic coding traffic: 98.2% cached input, 1.6% uncached input, 0.2% output. Rows without context compression carry 40% more tokens for the same work, because the harness re-sends a conversation that grows with every step. Our GLM-5.3 rate is the same list price Z.ai publishes, so that row is the cost of compression alone, not a cheaper model.

Enterprise

Deploy in your cloud

Past a certain scale it is cheaper to own the serving. We install our inference system inside your VPC, on your GPUs, and charge a monthly fee per node instead of per token.

For enterprises

Half the GPUs, faster throughput, better capability, better economics

Our serving infrastructure runs more concurrent agents on the same hardware, so a node does the work that used to take two. You keep the savings, we charge a fraction of it, with hands-on setup by our team of world-class AI researchers.

PricingMonthly fee per GPU node
DeploymentInstalled in your VPC
ModelsAny open model
SupportWorld-class AI expert support

50%

Same work with half the GPUs

2x

Faster task completion

5x+

Extended context window

Expert

Support from our team

If you have GPUs

We install in your cloud

Already running GPUs? We deploy our inference system directly onto the fleet you have today, and your team is serving open models in your own cloud within a week.

If you need GPUs

We get you the hardware first

No spare GPUs? We connect you to our compute partners or work with your cloud provider of choice, secure the lowest cost per GPU we can, then install and run the system for you.

Estimate your savings

GPU price, per hour

≈ H100
≈ B200 / B300

Deployment size, GPUs

Standard inference infrastructure

$186,880/mo

 

With Subconscious

$116,800/mo

50% the GPUs + our per GPU pricing*

You save

$70,080/mo

38% lower per month

* An estimate of our cost per node using our inference system, which covers everything we provide: our inference system, a routing gateway and supporting infrastructure, a frontend for API key management and usage monitoring, and dedicated setup and support from our team of world-class AI researchers.

Get started

Start on a plan today, in your cloud tomorrow.

Pick a monthly plan and point your harness at our API in a couple of minutes, or talk to us about installing the whole system on your own GPUs.