# xAI vs Inference.net

> A closed model lab against a platform that routes, captures and distills traffic into custom models. Grok is a model to call; Inference.net is a path off closed APIs.

Canonical: https://www.subconscious.dev/compare/xai-vs-inference-net · By The Subconscious Team · Updated September 30, 2026

## How they compare

Inference.net's product assumes you already use a closed model. Its Inference Gateway routes traffic to open, closed or custom models under one key, captures every request, and turns that traffic into eval and training datasets. It then fine-tunes a task-specific model and serves it on a dedicated GPU with a 99.99% uptime target. It also runs a Batch API on spare GPU capacity that takes up to 1M requests per file. xAI is the kind of closed provider that gateway might sit in front of, with Grok 4.6 at $2 in and $6 out and native X Search.

So the question is scope. Grok is the right call when the task needs live X data, strict prompt adherence or a general reasoning model that is cheap on output. Inference.net fits when a narrow, high-volume workload, like extraction or classification, can move to a smaller custom model to cut cost and latency. Its fragmented capacity suits batch better than strict real-time SLAs, and it has few independent benchmarks. xAI's doubled bill past 200K prompt tokens is another reason to push long bulk jobs elsewhere.

## What each one does

### xAI

xAI sells the Grok models through its own API. Grok 4.6 is the current flagship and xAI tells developers to use it for everything outside audio, image and video, code included. It has a 500K context window and costs $2 in and $6 out per million tokens under 200K prompt tokens. Older Grok 4.20 and 4.3 models keep a 1M window at $1.25 in and $2.50 out, which is aggressive for that capability class.

### Inference.net

Inference.net started as a buyer of last resort for idle GPU time. Its scheduler aggregates small unused chunks of capacity across data centers and runs models on them, and it passes the steep discounts it gets from those data centers on to customers. That origin still shows in its OpenAI-compatible Batch API, which takes up to 1M requests per file with completion windows from 24 hours to 7 days and far higher headroom than synchronous limits.

## Which is best, and when

### Choose xAI for

- General reasoning with live X and web search
- Interactive agents that need answers now
- Cheap output on a closed model

### Choose Inference.net for

- Distilling a narrow workload into a smaller custom model
- Large offline jobs on discounted spare capacity
- Capturing traffic for evals and training data

## At a glance

| Attribute | xAI | Inference.net |
|---|---|---|
| Model access | Closed | Open, closed and custom |
| Flagship models | Grok 4.6, Grok 4.20, grok-build | Customer fine-tunes |
| Speed | ~54 tok/s on Grok 4.6 | Batch windows of 24h to 7 days |
| Price | $2 in, $6 out (Grok 4.6); 2x past 200K | Discounted spare GPU capacity |
| Customization | - | Distill traces into custom models |
| Deployment | First-party API | Batch API, gateway, dedicated GPUs |
| Long context | 500K (4.6), 1M (4.20, 4.3) | Varies by model |

## FAQ

### What is the difference between xAI and Inference.net?

A closed model lab against a platform that routes, captures and distills traffic into custom models. Grok is a model to call; Inference.net is a path off closed APIs.

### When should I choose xAI over Inference.net?

General reasoning with live X and web search; Interactive agents that need answers now; Cheap output on a closed model.

### When should I choose Inference.net over xAI?

Distilling a narrow workload into a smaller custom model; Large offline jobs on discounted spare capacity; Capturing traffic for evals and training data.

### Is xAI or Inference.net cheaper?

xAI: $2 in, $6 out (Grok 4.6); 2x past 200K. Inference.net: Discounted spare GPU capacity. The cheaper choice depends on the model and workload.

### Which has more context, xAI or Inference.net?

xAI: 500K (4.6), 1M (4.20, 4.3). Inference.net: Varies by model.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs xAI](https://www.subconscious.dev/compare/subconscious-vs-xai.md), [Subconscious vs Inference.net](https://www.subconscious.dev/compare/subconscious-vs-inference-net.md).

Full profiles: [xAI](https://www.subconscious.dev/providers/xai.md), [Inference.net](https://www.subconscious.dev/providers/inference-net.md).
