Pricing
Usage based pricing or monthly plans to power your agents efficiently.
Monthly token plans
Tokens nearly too cheap to meter
Pick a daily allocation and share it across as many people as you want. No per-seat pricing, no per-request metering. Use it for coding or your product.
Base
Get moving fast on Subconscious
60Mtokens per day
- Every model on the API
- Drop-in for Claude Code, Codex, OpenCode, or any OpenAI-compatible harness
Cancel any time
Most popular
Pro
For a small team or heavy users
300Mtokens per day
- Everything in Base, with 5x the daily allocation
- Enough concurrency to fan out agents across a large codebase
Cancel any time
Heavy
For a team of heavy users
1.2Btokens per day
- Everything in Pro, with 4x the daily allocation
- Per-member API keys and a shared usage dashboard
- Direct line to our engineering team
Cancel any time
Build your own
Custom
Dial in the exact allocation your team burns through
$5,000/ mo
$200 to $10,000, in $100 steps
- Any allocation between the bounds, in $100 steps
- Per-member API keys and a shared usage dashboard
- Direct line to our engineering team
Cancel any time
Every plan includes
- Use it solo, or invite your team and share the daily allocation between you
- Hard daily ceiling by default, so a runaway agent cannot surprise you
- Add credits to keep going past the ceiling at the per-token rates below
- Switch tiers or cancel at any time from the dashboard
Ceilings reset daily, so the monthly figure assumes a 30 day month. Tokens per month assume a token mix of 98.2% cached, 1.6% input, and 0.2% output, characteristic of long-horizon tasks. Need more than 6B tokens a day, or a dedicated endpoint? Talk to us about a deployment in your cloud below.
Usage based pricing
Or pay only for the tokens you use
Pay per token with no plan and no commitment. Working past your token plan's ceiling? Tack on additional tokens at these standard rates.
| Model | Cached tokens | Input tokens | Output tokens |
|---|---|---|---|
| GLM-5.3 Marathon Open-source frontier coding model subconscious/glm-5.3-marathon | $0.26 | $1.40 | $4.40 |
| DeepSeek V4 Flash Marathon Multimodal & high-throughput DeepSeek V4 subconscious/deepseek-v4-flash-marathon | $0.0028 | $0.14 | $0.28 |
| Kimi K3 | Available for dedicated deploymentContact us | ||
| GLM-5.3 Flash | Available for dedicated deploymentContact us | ||
| DeepSeek V4 Pro | Available for dedicated deploymentContact us | ||
| Gemma 4 31B | Available for dedicated deploymentContact us | ||
| gpt-oss-120b | Available for dedicated deploymentContact us | ||
| Nemotron 3 Ultra | Available for dedicated deploymentContact us | ||
| Inkling | Available for dedicated deploymentContact us | ||
Prices in USD per 1M tokens. With efficient caching, we see upwards of a 95% cache hit rate on agentic coding workloads, so most input tokens bill at the cached rate.
What 9B tokens per month costs
Power your agents with more tokens, for a lot less
Subconscious Pro plan
$500Flat monthly fee, shared across your team
Subconscious, pay per token
$2,579Billed per request
GLM-5.3 on other platforms
$3,610Identical list price, but without context compression
Claude Opus 5
$7,825Frontier list price
Assumes agentic coding traffic: 98.2% cached input, 1.6% uncached input, 0.2% output. Rows without context compression carry 40% more tokens for the same work, because the harness re-sends a conversation that grows with every step. Our GLM-5.3 rate is the same list price Z.ai publishes, so that row is the cost of compression alone, not a cheaper model.
Enterprise
Deploy in your cloud
Past a certain scale it is cheaper to own the serving. We install our inference system inside your VPC, on your GPUs, and charge a monthly fee per node instead of per token.
For enterprises
Half the GPUs, faster throughput, better capability, better economics
Our serving infrastructure runs more concurrent agents on the same hardware, so a node does the work that used to take two. You keep the savings, we charge a fraction of it, with hands-on setup by our team of world-class AI researchers.
50%
Same work with half the GPUs
2x
Faster task completion
5x+
Extended context window
Expert
Support from our team
If you have GPUs
We install in your cloud
Already running GPUs? We deploy our inference system directly onto the fleet you have today, and your team is serving open models in your own cloud within a week.
If you need GPUs
We get you the hardware first
No spare GPUs? We connect you to our compute partners or work with your cloud provider of choice, secure the lowest cost per GPU we can, then install and run the system for you.
Estimate your savings
GPU price, per hour
Deployment size, GPUs
Standard inference infrastructure
$186,880/mo
With Subconscious
$116,800/mo
50% the GPUs + our per GPU pricing*
You save
$70,080/mo
38% lower per month
* An estimate of our cost per node using our inference system, which covers everything we provide: our inference system, a routing gateway and supporting infrastructure, a frontend for API key management and usage monitoring, and dedicated setup and support from our team of world-class AI researchers.