ROI calculator
Model the GPU economics behind coding agents.
Estimate capacity, concurrency, and monthly savings when you move high-volume coding-agent inference from API spend to managed GPUs running the Subconscious runtime.
Share by URL
Every calculator assumption lives in the link. Copy it to share a scenario or bookmark it to reopen the same model later.
Calculator workspace
Tune assumptions and compare scenarios.
Tokens / engineer / day
100M
Normalized from selected period
Tokens / day
10B
Engineer demand
Service level
Demand varies within the hour even when the daily total is fixed. Capacity is sized to cover that variation this share of the time.
Peak mean TPS
83.5K tokens/s
Peak TPS (P95)
279.7K tokens/s
Mean 83.5K · 235.0% margin
Hourly bars use OrangeLine GPU-processed demand (vendor totals from Coding-agent demand are scaled by the trace-model ratio).
0.5625 bytes / parameter · 424 GB across the replica · 106.0 GB / GPU
54.0 GB of KV per GPU after 106.0 GB of sharded weights, so 8.5M tokens per replica. A min(window, peak) trace holds 150K resident tokens (3.8 GB), so memory admits 48 traces per replica. The concurrency reference of 32 is the fair-share measurement point, not an admission cap.
Agent trace
Output share: 16.0%
Simulated mix: 0.6% uncached · 99.2% cached · 0.1% output
Standard agent
200K summary tokens
OrangeLine pruning
Context size by turn
Traditional tokens / trace
67.8M
OrangeLine tokens / trace
29.4M
Trace token reduction
56.6%
Compaction events
0
Context extension
1x
Traditional KV / trace
13.21 GB
OrangeLine KV / trace
3.81 GB
KV memory reduction
71.2%
Tokens per trace, traditional vs OrangeLine
KV cache per trace, traditional vs OrangeLine
OrangeLine
Traces / replica by KV memory
48
150K context tokens · 3.8 GB / trace
Traditional
Traces / replica by KV memory
13
520K context tokens · 13.2 GB / trace
OrangeLine
Derived replica throughput
64.1K tokens/s
Average trace duration
1.1 hr
Traditional
Derived replica throughput
87K tokens/s
Average trace duration
1.8 hr
OrangeLine
Tokens / day
4.3B
P95 throughput demand
279.7K tok/s
Avg trace duration
1.1 hr
Recommended replicas
5
20 GPUs
Traditional
Tokens / day
10B
P95 throughput demand
644.3K tok/s
Avg trace duration
1.8 hr
Recommended replicas
8
32 GPUs
P95 throughput demand vs capacity
P95 token demand for both fleets against provisioned capacity.
P95 concurrency demand vs capacity
P95 concurrent traces for both fleets against provisioned slot capacity.
Next step
Want us to validate the model against your workload?
Bring your traces, token mix, and GPU assumptions. We can help translate this model into a deployment plan.