# Nebius vs Inference.net

> A European cloud with SLA-backed endpoints against a batch specialist built on spare GPU capacity that also distills traces into custom models.

Canonical: https://www.subconscious.dev/compare/nebius-vs-inference-net · By The Subconscious Team · Updated September 30, 2026

## How they compare

Nebius and Inference.net sit at opposite ends of capacity planning. Nebius runs a full AI cloud with its own GPUs, real-time Token Factory endpoints across 60+ open models, and dedicated deployments with a 99.9% SLA and speculative decoding. Inference.net began as a buyer of idle GPU time, and its OpenAI-compatible Batch API still reflects that: up to 1M requests per file, completion windows from 24 hours to 7 days, and pricing built on discounted spare capacity. That fragmented supply suits batch work better than strict real-time SLAs.

The second axis is customization. Both let a team run a fine-tuned model, but they approach it differently. Nebius serves a checkpoint you upload at the same token pricing as the base. Inference.net runs the whole loop for you: its gateway captures production traffic, turns it into eval and training data, fine-tunes a task-specific model and deploys it on a dedicated GPU with a 99.99% uptime target. Pick Nebius for interactive traffic, EU placement and training on your own terms. Pick Inference.net for large offline jobs or for replacing a narrow closed-model workload with a smaller distilled one. Buyers should note that Inference.net has few independent benchmarks.

## What each one does

### Nebius

Nebius is an Amsterdam-headquartered AI cloud and the strongest European alternative to the US hyperscalers. It sells raw NVIDIA GPU compute, from H100s at $2.15 an hour preemptible up to GB300 NVL72 racks, and it has begun adding Vera Rubin. Hyperscale buyers back it: a Microsoft capacity deal worth about $17.4B in September 2025, then a Meta agreement worth up to about $27B in March 2026.

### Inference.net

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

## Which is best, and when

### Choose Nebius for

- Real-time inference with a 99.9% SLA and EU or US placement
- Serving checkpoints your own team fine-tuned
- Measured high throughput on shared open models

### Choose Inference.net for

- Million-request batch files that can wait 24 hours or more
- Turning production traces into a distilled task-specific model
- Routing open, closed and custom models under one gateway key

## At a glance

| Attribute | Nebius | Inference.net |
|---|---|---|
| Model access | Open weights, 60+ models | Open, closed and custom |
| Flagship models | DeepSeek, Qwen, GLM, Kimi, GPT-OSS | Customer fine-tunes |
| Speed | Among top hosts on throughput | Batch windows of 24h to 7 days |
| Price | From $0.06 per 1M input | Discounted spare GPU capacity |
| Customization | Serve uploaded fine-tunes | Distill traces into custom models |
| Deployment | Token Factory, dedicated, raw GPUs | Batch API, gateway, dedicated GPUs |
| Long context | Varies by model | Varies by model |

## FAQ

### What is the difference between Nebius and Inference.net?

A European cloud with SLA-backed endpoints against a batch specialist built on spare GPU capacity that also distills traces into custom models.

### When should I choose Nebius over Inference.net?

Real-time inference with a 99.9% SLA and EU or US placement; Serving checkpoints your own team fine-tuned; Measured high throughput on shared open models.

### When should I choose Inference.net over Nebius?

Million-request batch files that can wait 24 hours or more; Turning production traces into a distilled task-specific model; Routing open, closed and custom models under one gateway key.

### Is Nebius or Inference.net cheaper?

Nebius: From $0.06 per 1M input. Inference.net: Discounted spare GPU capacity. The cheaper choice depends on the model and workload.

### Which has more context, Nebius or Inference.net?

Nebius: Varies by model. Inference.net: Varies by model.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Nebius](https://www.subconscious.dev/compare/subconscious-vs-nebius.md), [Subconscious vs Inference.net](https://www.subconscious.dev/compare/subconscious-vs-inference-net.md).

Full profiles: [Nebius](https://www.subconscious.dev/providers/nebius.md), [Inference.net](https://www.subconscious.dev/providers/inference-net.md).
