# Subconscious pricing

> Source of truth: https://www.subconscious.dev/pricing. All prices in USD.

Three ways to pay: usage based (per token), monthly token plans, or an enterprise deployment of the inference system inside your own cloud.

## Usage based pricing

Pay per token with no plan and no commitment. Prices per 1M tokens.

| Model | API model string | Cached input | Input | Output |
|---|---|---|---|---|
| GLM-5.3 Marathon | `subconscious/glm-5.3-marathon` | $0.26 | $1.40 | $4.40 |
| DeepSeek V4.1 Flash Marathon | `subconscious/deepseek-v4-flash-marathon` | $0.0028 | $0.14 | $0.28 |

Cached input tokens bill at a steep discount. With efficient caching we see upwards of a 95% cache hit rate on agentic coding workloads, so most input tokens bill at the cached rate.

## Monthly token plans

Fixed monthly price for a daily token allowance, shared across models and across your team.

| Plan | Price | Tokens per day | Tokens per month | Concurrent requests | For |
|---|---|---|---|---|---|
| Base | $100/month | 60M | 1.8B | 10 | Get moving fast on Subconscious |
| Pro | $500/month | 300M | 9B | 50 | For a small team or heavy users |
| Heavy | $2,000/month | 1.2B | 36B | 200 | For a team of heavy users |

- Custom: any allocation from $200 to $10,000 a month in $100 steps. Each step adds 60M tokens per day and 10 concurrent requests.
- Use it solo, or invite your team and share the daily allocation between you
- Hard daily ceiling by default, so a runaway agent cannot surprise you
- Add credits to keep going past the ceiling at the per-token rates below
- Switch tiers or cancel at any time from the dashboard
- Daily allocations reset at midnight UTC. Allocation figures assume an agentic token mix of 98.2% cached input, 1.6% uncached input, 0.2% output.

## Enterprise

Run the Subconscious inference system on your own GPUs, installed and supported by our team, 100% inside your cloud. Priced per GPU. Dedicated endpoints are also available.

- Contact: https://www.subconscious.dev/enterprise/contact
