# Inference.net

> Cheap batch inference on spare GPU capacity, plus a path from traces to custom models.

Canonical: https://www.subconscious.dev/providers/inference-net · By The Subconscious Team · Updated September 30, 2026

- Founded: 2023
- Example models: open catalog models plus customer fine-tunes served on dedicated GPUs
- Website: https://inference.net

## Overview

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

The company now pitches a full loop for teams that want to leave closed APIs. Inference Gateway routes traffic to open, closed or custom models under one key, captures every request, and turns that traffic into eval and training datasets. From there, Inference.net fine-tunes a task-specific model and deploys it on a dedicated GPU with a 99.99% uptime target. A tracing product with an open-source optimizer called Halo reads agent traces and suggests fixes to prompts, tools and harness. Customers named on the site include GravityAds and Cal AI.

## Upsides

- Cheap bulk inference built on otherwise wasted GPU capacity.
- One workflow from production traces to a distilled custom model.

## Core use cases

- Large offline jobs like extraction, classification and synthetic data.
- Replacing a narrow GPT-class workload with a smaller fine-tuned model to cut cost and latency.

## Downsides

- Fragmented spare capacity suits batch work better than strict real-time SLAs.
- Few independent benchmarks or public pricing comparisons, so buyers depend on the vendor's numbers.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open, closed and custom |
| Flagship models | Customer fine-tunes |
| Speed | Batch windows of 24h to 7 days |
| Price | Discounted spare GPU capacity |
| Customization | Distill traces into custom models |
| Deployment | Batch API, gateway, dedicated GPUs |
| Long context | Varies by model |

## FAQ

### What is Inference.net?

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

### What is Inference.net best for?

Large offline jobs like extraction, classification and synthetic data; Replacing a narrow GPT-class workload with a smaller fine-tuned model to cut cost and latency.

### How much does Inference.net cost?

Inference.net pricing at a glance: Discounted spare GPU capacity. Rates change often, so check Inference.net's pricing page before committing.

### How much context does Inference.net support?

Inference.net's long-context support: Varies by model.

### What are the downsides of Inference.net?

Fragmented spare capacity suits batch work better than strict real-time SLAs; Few independent benchmarks or public pricing comparisons, so buyers depend on the vendor's numbers.

### What are the best alternatives to Inference.net?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Inference.net on this site.

## Comparisons

- [Subconscious vs Inference.net](https://www.subconscious.dev/compare/subconscious-vs-inference-net.md)
- [OpenAI vs Inference.net](https://www.subconscious.dev/compare/openai-vs-inference-net.md)
- [Anthropic vs Inference.net](https://www.subconscious.dev/compare/anthropic-vs-inference-net.md)
- [Google Vertex AI vs Inference.net](https://www.subconscious.dev/compare/google-vertex-vs-inference-net.md)
- [Amazon Bedrock vs Inference.net](https://www.subconscious.dev/compare/aws-bedrock-vs-inference-net.md)
- [Together AI vs Inference.net](https://www.subconscious.dev/compare/together-ai-vs-inference-net.md)
- [Fireworks AI vs Inference.net](https://www.subconscious.dev/compare/fireworks-vs-inference-net.md)
- [Baseten vs Inference.net](https://www.subconscious.dev/compare/baseten-vs-inference-net.md)
- [Groq vs Inference.net](https://www.subconscious.dev/compare/groq-vs-inference-net.md)
- [Cerebras vs Inference.net](https://www.subconscious.dev/compare/cerebras-vs-inference-net.md)
- [DeepInfra vs Inference.net](https://www.subconscious.dev/compare/deepinfra-vs-inference-net.md)
- [Modal vs Inference.net](https://www.subconscious.dev/compare/modal-vs-inference-net.md)
- [xAI vs Inference.net](https://www.subconscious.dev/compare/xai-vs-inference-net.md)
- [DeepSeek vs Inference.net](https://www.subconscious.dev/compare/deepseek-vs-inference-net.md)
- [Moonshot AI vs Inference.net](https://www.subconscious.dev/compare/moonshot-ai-vs-inference-net.md)
- [Z.ai vs Inference.net](https://www.subconscious.dev/compare/z-ai-vs-inference-net.md)
- [Alibaba Cloud vs Inference.net](https://www.subconscious.dev/compare/alibaba-cloud-vs-inference-net.md)
- [Meta vs Inference.net](https://www.subconscious.dev/compare/meta-vs-inference-net.md)
- [SambaNova vs Inference.net](https://www.subconscious.dev/compare/sambanova-vs-inference-net.md)
- [Nebius vs Inference.net](https://www.subconscious.dev/compare/nebius-vs-inference-net.md)
- [fal vs Inference.net](https://www.subconscious.dev/compare/fal-vs-inference-net.md)
- [Novita AI vs Inference.net](https://www.subconscious.dev/compare/novita-ai-vs-inference-net.md)
- [Parasail vs Inference.net](https://www.subconscious.dev/compare/parasail-vs-inference-net.md)
- [Inference.net vs GMI Cloud](https://www.subconscious.dev/compare/inference-net-vs-gmi-cloud.md)
- [Inference.net vs Sail Research](https://www.subconscious.dev/compare/inference-net-vs-sail-research.md)
- [Inference.net vs Morph](https://www.subconscious.dev/compare/inference-net-vs-morph.md)
- [Inference.net vs Relace](https://www.subconscious.dev/compare/inference-net-vs-relace.md)
- [Inference.net vs TypeSafe AI](https://www.subconscious.dev/compare/inference-net-vs-typesafe-ai.md)
- [Inference.net vs StepFun](https://www.subconscious.dev/compare/inference-net-vs-stepfun.md)
- [Inference.net vs Runware](https://www.subconscious.dev/compare/inference-net-vs-runware.md)
- [Inference.net vs StreamLake](https://www.subconscious.dev/compare/inference-net-vs-streamlake.md)
- [Inference.net vs Wafer](https://www.subconscious.dev/compare/inference-net-vs-wafer.md)
- [Inference.net vs RunInfra](https://www.subconscious.dev/compare/inference-net-vs-runinfra.md)
- [Inference.net vs Particle.AI](https://www.subconscious.dev/compare/inference-net-vs-particle-ai.md)

## Sources

- [Inference.net](https://inference.net/)
- [Inference.net company](https://inference.net/company/)
- [Batch API docs](https://docs.inference.net/api/async-inference/batch-api)

Pricing and model lineups change often; figures are a snapshot.
