# Hugging Face Inference Providers

> One Hugging Face token and one OpenAI-compatible router in front of 17 partner inference providers, at provider rates.

Canonical: https://www.subconscious.dev/providers/hugging-face · By The Subconscious Team · Updated September 30, 2026

- Founded: 2016
- Example models: GLM 5.3, Kimi K3, gpt-oss-120b
- Website: https://huggingface.co/docs/inference-providers

## Overview

Inference Providers is a router run by Hugging Face that sits in front of partner inference clouds. The current partner list covers Baseten, Cerebras, Cohere, DeepInfra, fal, Featherless AI, Fireworks, Groq, Novita, Nscale, OVHcloud, Public AI, Replicate, Scaleway, Together, WaveSpeedAI and Z.ai, plus Hugging Face's own HF Inference, which now mostly serves CPU workloads like embeddings and classification. Chat traffic goes through an OpenAI-compatible endpoint at router.huggingface.co/v1, and the Python and JavaScript clients add text-to-image, video, speech and embeddings. The router lists 132 chat models today, from GLM 5.3 and Kimi K3 to gpt-oss-120b on eleven providers.

By default the router picks the provider with the highest throughput for a model. Appending :cheapest to the model id picks the lowest price per output token, :preferred follows an order set in account settings, and a provider name like :groq pins one host. Automatic selection fails over when a provider is flagged unavailable. Hugging Face bills at the provider's rate with no markup, or users can bring their own provider key and be billed directly. Free accounts get $0.10 a month in credits and PRO users $2. For contrast, Inference Endpoints is the dedicated option: per-minute billing on AWS, GCP or Azure with vLLM, SGLang, TGI or llama.cpp, from $0.50 an hour for a T4.

## Upsides

- One token, one bill and one OpenAI-compatible API across 17 partner providers.
- Pass-through pricing with no markup, and :fastest or :cheapest routing without code changes.
- Live per-provider price, context, latency and throughput exposed through /v1/models.

## Core use cases

- Comparing the same open model across hosts before committing to one.
- Prototyping on new open models straight from their Hub model pages.
- Consolidating inference spend for a team under one organization bill.

## Downsides

- Adds a network hop and Hugging Face rate limits on top of each provider's own.
- No fine-tuning, and the OpenAI-compatible endpoint covers chat only, so other tasks need the Hugging Face SDKs.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open weights |
| Flagship models | GLM 5.3, Kimi K3, DeepSeek V4.1 Flash |
| Speed | Routes to fastest provider by default |
| Price | Provider rates, no markup |
| Customization | N/A |
| Deployment | Serverless router; dedicated Endpoints |
| Long context | Up to 1M, provider-dependent |

## FAQ

### What is Hugging Face Inference Providers?

Inference Providers is a router run by Hugging Face that sits in front of partner inference clouds. The current partner list covers Baseten, Cerebras, Cohere, DeepInfra, fal, Featherless AI, Fireworks, Groq, Novita, Nscale, OVHcloud, Public AI, Replicate, Scaleway, Together, WaveSpeedAI and Z.ai, plus Hugging Face's own HF Inference, which now mostly serves CPU workloads like embeddings and classification. Chat traffic goes through an OpenAI-compatible endpoint at router.huggingface.co/v1, and the Python and JavaScript clients add text-to-image, video, speech and embeddings. The router lists 132 chat models today, from GLM 5.3 and Kimi K3 to gpt-oss-120b on eleven providers.

### What is Hugging Face Inference Providers best for?

Comparing the same open model across hosts before committing to one; Prototyping on new open models straight from their Hub model pages; Consolidating inference spend for a team under one organization bill.

### How much does Hugging Face Inference Providers cost?

Hugging Face Inference Providers pricing at a glance: Provider rates, no markup. Rates change often, so check Hugging Face Inference Providers's pricing page before committing.

### How much context does Hugging Face Inference Providers support?

Hugging Face Inference Providers's long-context support: Up to 1M, provider-dependent.

### What are the downsides of Hugging Face Inference Providers?

Adds a network hop and Hugging Face rate limits on top of each provider's own; No fine-tuning, and the OpenAI-compatible endpoint covers chat only, so other tasks need the Hugging Face SDKs.

### What are the best alternatives to Hugging Face Inference Providers?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Hugging Face Inference Providers on this site.

## Comparisons

- [Subconscious vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/subconscious-vs-hugging-face.md)
- [OpenAI vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/openai-vs-hugging-face.md)
- [Anthropic vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/anthropic-vs-hugging-face.md)
- [Google Vertex AI vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/google-vertex-vs-hugging-face.md)
- [Amazon Bedrock vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/aws-bedrock-vs-hugging-face.md)
- [Together AI vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/together-ai-vs-hugging-face.md)
- [Fireworks AI vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/fireworks-vs-hugging-face.md)
- [Baseten vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/baseten-vs-hugging-face.md)
- [Groq vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/groq-vs-hugging-face.md)
- [Cerebras vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/cerebras-vs-hugging-face.md)
- [DeepInfra vs Hugging Face Inference Providers](https://www.subconscious.dev/compare/deepinfra-vs-hugging-face.md)
- [Hugging Face Inference Providers vs Modal](https://www.subconscious.dev/compare/hugging-face-vs-modal.md)
- [Hugging Face Inference Providers vs Cloudflare Workers AI](https://www.subconscious.dev/compare/hugging-face-vs-cloudflare-workers-ai.md)
- [Hugging Face Inference Providers vs xAI](https://www.subconscious.dev/compare/hugging-face-vs-xai.md)
- [Hugging Face Inference Providers vs Mistral AI](https://www.subconscious.dev/compare/hugging-face-vs-mistral-ai.md)
- [Hugging Face Inference Providers vs DeepSeek](https://www.subconscious.dev/compare/hugging-face-vs-deepseek.md)
- [Hugging Face Inference Providers vs Moonshot AI](https://www.subconscious.dev/compare/hugging-face-vs-moonshot-ai.md)
- [Hugging Face Inference Providers vs Z.ai](https://www.subconscious.dev/compare/hugging-face-vs-z-ai.md)
- [Hugging Face Inference Providers vs Alibaba Cloud](https://www.subconscious.dev/compare/hugging-face-vs-alibaba-cloud.md)
- [Hugging Face Inference Providers vs Meta](https://www.subconscious.dev/compare/hugging-face-vs-meta.md)
- [Hugging Face Inference Providers vs Cohere](https://www.subconscious.dev/compare/hugging-face-vs-cohere.md)
- [Hugging Face Inference Providers vs SambaNova](https://www.subconscious.dev/compare/hugging-face-vs-sambanova.md)
- [Hugging Face Inference Providers vs Nebius](https://www.subconscious.dev/compare/hugging-face-vs-nebius.md)
- [Hugging Face Inference Providers vs Crusoe](https://www.subconscious.dev/compare/hugging-face-vs-crusoe.md)
- [Hugging Face Inference Providers vs fal](https://www.subconscious.dev/compare/hugging-face-vs-fal.md)
- [Hugging Face Inference Providers vs Novita AI](https://www.subconscious.dev/compare/hugging-face-vs-novita-ai.md)
- [Hugging Face Inference Providers vs Venice](https://www.subconscious.dev/compare/hugging-face-vs-venice.md)
- [Hugging Face Inference Providers vs Parasail](https://www.subconscious.dev/compare/hugging-face-vs-parasail.md)
- [Hugging Face Inference Providers vs Inference.net](https://www.subconscious.dev/compare/hugging-face-vs-inference-net.md)
- [Hugging Face Inference Providers vs GMI Cloud](https://www.subconscious.dev/compare/hugging-face-vs-gmi-cloud.md)
- [Hugging Face Inference Providers vs Thinking Machines](https://www.subconscious.dev/compare/hugging-face-vs-thinking-machines.md)
- [Hugging Face Inference Providers vs Sail Research](https://www.subconscious.dev/compare/hugging-face-vs-sail-research.md)
- [Hugging Face Inference Providers vs Morph](https://www.subconscious.dev/compare/hugging-face-vs-morph.md)
- [Hugging Face Inference Providers vs Relace](https://www.subconscious.dev/compare/hugging-face-vs-relace.md)
- [Hugging Face Inference Providers vs TypeSafe AI](https://www.subconscious.dev/compare/hugging-face-vs-typesafe-ai.md)
- [Hugging Face Inference Providers vs StepFun](https://www.subconscious.dev/compare/hugging-face-vs-stepfun.md)
- [Hugging Face Inference Providers vs Runware](https://www.subconscious.dev/compare/hugging-face-vs-runware.md)
- [Hugging Face Inference Providers vs StreamLake](https://www.subconscious.dev/compare/hugging-face-vs-streamlake.md)
- [Hugging Face Inference Providers vs Wafer](https://www.subconscious.dev/compare/hugging-face-vs-wafer.md)
- [Hugging Face Inference Providers vs RunInfra](https://www.subconscious.dev/compare/hugging-face-vs-runinfra.md)
- [Hugging Face Inference Providers vs Particle.AI](https://www.subconscious.dev/compare/hugging-face-vs-particle-ai.md)

## Sources

- [Inference Providers docs](https://huggingface.co/docs/inference-providers/index)
- [Inference Providers pricing](https://huggingface.co/docs/inference-providers/pricing)
- [Inference Endpoints pricing](https://huggingface.co/docs/inference-endpoints/pricing)
- [Hugging Face vs Fireworks, Markaicode](https://markaicode.com/vs/hugging-face-vs-fireworks-ai/)

Pricing and model lineups change often; figures are a snapshot.
