ROI calculator

Model the GPU economics behind coding agents.

Estimate capacity, concurrency, and monthly savings when you move high-volume coding-agent inference from API spend to managed GPUs running the Subconscious runtime.

Share by URL

Every calculator assumption lives in the link. Copy it to share a scenario or bookmark it to reopen the same model later.

Calculator workspace

Tune assumptions and compare scenarios.

Tokens / engineer / day

100M

Normalized from selected period

Tokens / day

10B

Engineer demand

Service level

Demand varies within the hour even when the daily total is fixed. Capacity is sized to cover that variation this share of the time.

90%

Peak mean TPS

83.5K tokens/s

Peak TPS (P95)

279.7K tokens/s

Mean 83.5K · 235.0% margin

Hourly bars use OrangeLine GPU-processed demand (vendor totals from Coding-agent demand are scaled by the trace-model ratio).

GB
B

0.5625 bytes / parameter · 424 GB across the replica · 106.0 GB / GPU

GB
90%
85%

54.0 GB of KV per GPU after 106.0 GB of sharded weights, so 8.5M tokens per replica. A min(window, peak) trace holds 150K resident tokens (3.8 GB), so memory admits 48 traces per replica. The concurrency reference of 32 is the fair-share measurement point, not an admission cap.

Agent trace

Turns per trace250
Tokens / turn2,000
Generated input share / turn84%

Output share: 16.0%

Simulated mix: 0.6% uncached · 99.2% cached · 0.1% output

Standard agent

Compaction target20%

200K summary tokens

OrangeLine pruning

Context size by turn
OrangeLine contextTraditional contextEffective max context

Traditional tokens / trace

67.8M

OrangeLine tokens / trace

29.4M

Trace token reduction

56.6%

Compaction events

0

Context extension

1x

Traditional KV / trace

13.21 GB

OrangeLine KV / trace

3.81 GB

KV memory reduction

71.2%

Tokens per trace, traditional vs OrangeLine
Uncached inputCached inputNormal outputCompaction output
KV cache per trace, traditional vs OrangeLine
Traditional KVOrangeLine KV

OrangeLine

Traces / replica by KV memory

48

150K context tokens · 3.8 GB / trace

Traditional

Traces / replica by KV memory

13

520K context tokens · 13.2 GB / trace

OrangeLine

Derived replica throughput

64.1K tokens/s

Average trace duration

1.1 hr

Traditional

Derived replica throughput

87K tokens/s

Average trace duration

1.8 hr

OrangeLine

Tokens / day

4.3B

P95 throughput demand

279.7K tok/s

Avg trace duration

1.1 hr

Recommended replicas

5

20 GPUs

Traditional

Tokens / day

10B

P95 throughput demand

644.3K tok/s

Avg trace duration

1.8 hr

Recommended replicas

8

32 GPUs

P95 throughput demand vs capacity

P95 token demand for both fleets against provisioned capacity.

OrangeLine P95Traditional P95OrangeLine capacityTraditional capacity

P95 concurrency demand vs capacity

P95 concurrent traces for both fleets against provisioned slot capacity.

OrangeLine P95Traditional P95OrangeLine capacityTraditional capacity

Next step

Want us to validate the model against your workload?

Bring your traces, token mix, and GPU assumptions. We can help translate this model into a deployment plan.

Talk to us

© 2026 Subconscious Systems Technologies, Inc.

Subconscious