# Subconscious vs Inference.net

> Inference.net turns spare GPU capacity into cheap batch jobs. Subconscious powers the live, long-running agents that cannot wait a day for an answer.

Canonical: https://www.subconscious.dev/compare/subconscious-vs-inference-net · By The Subconscious Team · Updated September 30, 2026

## How they compare

Inference.net is built for work that can wait. Its Batch API takes up to 1M requests per file with completion windows from 24 hours to 7 days, running on spare GPU capacity bought at a steep discount. It also offers a gateway that captures traffic and turns it into datasets, then fine-tunes and deploys a task-specific model on a dedicated GPU. Subconscious is built for work that cannot wait and does not shrink into a small model: long-horizon agents whose traces pass 200K tokens. It prunes the KV cache, bills processed tokens, and delivers neutral to 10% better scores on agentic benchmarks.

Inference.net is the better pick for large offline extraction, classification and synthetic data, and for replacing a narrow GPT-class task with a smaller fine-tuned model. Its spare-capacity design suits batch better than strict real-time SLAs, and buyers depend mostly on its own numbers. For live agents, Subconscious adds 2x faster task completion and a 5M+ effective context window. The two can coexist. Inference.net's Halo optimizer reads agent traces and suggests prompt, tool and harness fixes, and it could review traces from an agent running on Subconscious.

## What each one does

### Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### Inference.net

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

## Which is best, and when

### Choose Subconscious for

- Live long-horizon agents that need answers inside the loop
- Tasks too broad to distill into a small model
- Traces past 200K tokens billed after compression

### Choose Inference.net for

- Offline jobs with 24-hour to 7-day windows
- Distilling production traffic into a small custom model
- Routing open, closed and custom models under one key

## At a glance

| Attribute | Subconscious | Inference.net |
|---|---|---|
| Model access | Open weights | Open, closed and custom |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | Customer fine-tunes |
| Speed | 2x faster task completion | Batch windows of 24h to 7 days |
| Price | 50–80% lower cost; billed on processed tokens | Discounted spare GPU capacity |
| Customization | Marathon post-trained variants | Distill traces into custom models |
| Deployment | Managed API, dedicated, on-prem | Batch API, gateway, dedicated GPUs |
| Long context | 5M+ effective context | Varies by model |

## FAQ

### What is the difference between Subconscious and Inference.net?

Inference.net turns spare GPU capacity into cheap batch jobs. Subconscious powers the live, long-running agents that cannot wait a day for an answer.

### When should I choose Subconscious over Inference.net?

Live long-horizon agents that need answers inside the loop; Tasks too broad to distill into a small model; Traces past 200K tokens billed after compression.

### When should I choose Inference.net over Subconscious?

Offline jobs with 24-hour to 7-day windows; Distilling production traffic into a small custom model; Routing open, closed and custom models under one key.

### Is Subconscious or Inference.net cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Inference.net: Discounted spare GPU capacity. The cheaper choice depends on the model and workload.

### Which has more context, Subconscious or Inference.net?

Subconscious: 5M+ effective context. Inference.net: Varies by model.

## Try Subconscious

Subconscious speaks the OpenAI and Anthropic API formats. Base URL: https://api.subconscious.dev/v1. Docs: https://docs.subconscious.dev. Get an API key: https://platform.subconscious.dev/signin. Agent guide: https://www.subconscious.dev/agents.md.

Full profiles: [Subconscious](https://www.subconscious.dev/providers/subconscious.md), [Inference.net](https://www.subconscious.dev/providers/inference-net.md).
