Every major inference provider, compared.
What each provider does, where it wins, and how it stacks up on the dimensions that decide a production workload.
OpenAI
The most widely adopted closed-model API, from GPT-6 Astra down to GPT-5.6 Luna.
Anthropic
Closed Claude models known for agentic coding, with 1M context and no long-context premium.
Google Vertex AI
Google Cloud's enterprise AI platform, with Gemini and Claude next to a full MLOps stack.
Amazon Bedrock
AWS's managed model service: Claude, GPT and open models inside your AWS security posture.
Together AI
The broadest open-model platform: serverless, dedicated, GPU clusters and fine-tuning on one bill.
Fireworks AI
Fast open-model inference plus SFT, DPO and reinforcement fine-tuning at base-model prices.
Baseten
Lowest measured time to first token, with OpenAI and Anthropic-compatible endpoints.
Groq
Extreme speed on its own LPU chip, with tight tail latency on a small catalog.
Cerebras
The fastest public inference host, running models on a wafer-scale chip.
DeepInfra
The price floor for open-model inference, with 150+ models and no minimums.
Modal
Serverless GPU compute for Python: bring your own model and pay by the second.
xAI
Closed Grok models with native live data from X and cheap output tokens.
DeepSeek
MIT-licensed open models at some of the lowest first-party prices anywhere.
Moonshot AI
The lab behind Kimi K3, the most capable open-weight model, with 1M context.
Z.ai
Maker of the MIT-licensed GLM models, with a cheap flat-rate coding plan.
Alibaba Cloud
Home of the Qwen family, with a closed Max flagship inside a full public cloud.
Meta
The Muse model family on a new closed API, plus open-weight Muse Glimmer.
SambaNova
Fast decode on large open models, served on its own dataflow chip.
Nebius
A European AI cloud with managed open-model inference and raw GPUs on one account.
fal
The go-to platform for generative media, with 1,000+ image, video and audio models.
Novita AI
A low-cost inference cloud with 200+ open models across text, image, video and speech.
Parasail
Aggregated GPUs behind one API, with cheap batch on any Hugging Face model.
Inference.net
Cheap batch inference on spare GPU capacity, plus a path from traces to custom models.
GMI Cloud
A GPU cloud that owns its hardware, with APAC data residency and 100+ models.
Sail Research
Slow inference for very cheap: pick a completion window and save 30 to 80%.
Morph
Small specialist models that apply coding-agent edits at 10,500+ tokens per second.
Relace
Fast small models that act as tools for coding agents: apply, search and compaction.

TypeSafe AI
Decision models that return typed answers with calibrated confidence in about 100ms.
StepFun
A Shanghai lab shipping efficient multimodal models, many under Apache 2.0.
Runware
Low-cost media generation across image, video, audio and 3D behind one request schema.
StreamLake
Kuaishou's AI cloud, home of the KAT-Coder agentic coding models.
Wafer
Agent-tuned inference stacks that run open models faster on the same weights.
RunInfra
Open models for agents, plus an agent that benchmarks and builds deployments for you.

Particle.AI
An early startup serving cheap Flash-class open models with 1M context.
Frequently asked questions
What is an AI inference provider?
An inference provider runs AI models for you and sells access by the token, the GPU hour or the request. Some are labs selling their own closed models (OpenAI, Anthropic, xAI), some are clouds hosting open-weight models (Together AI, Fireworks AI, Baseten, DeepInfra), some run custom chips for speed (Groq, Cerebras, SambaNova), and some are hyperscaler platforms (Amazon Bedrock, Google Vertex AI).
Which inference provider is the fastest?
Cerebras posts the fastest published output speed of any public host, near 3,000 tokens per second on GPT-OSS 120B. Groq serves 500 to 1,000 tokens per second with tight tail latency, and Baseten posted the lowest measured time to first token among the major open-model hosts at 0.49 seconds.
Which inference provider is the cheapest?
DeepInfra is the price floor for open-model inference, with small models from $0.02 per million tokens. Sail Research is cheaper still for work that can wait, at 30 to 80% off in exchange for minutes-long turns, and DeepSeek’s own API halves its prices off-peak.
Which inference provider is best for long-context and long-horizon agents?
Subconscious is built for agent traces past 200K tokens. It prunes the KV cache instead of rereading context, and against standard open-model inference it delivers 2x faster task completion, a 5M+ effective context window, and 50% to 80% lower cost. Anthropic offers a 1M window with no surcharge past 200K, while OpenAI and xAI charge more once prompts pass 272K and 200K tokens.
Which inference provider is best for fine-tuning open models?
Together AI covers LoRA and full-parameter SFT with reinforcement learning in beta, and serves the checkpoint on the same platform. Fireworks AI offers SFT, DPO and reinforcement fine-tuning, and serves fine-tuned models at the base model’s per-token price.
Which inference provider is best for image and video generation?
fal hosts 1,000+ image, video and audio models behind one API with billing that skips failures and cold starts. Runware competes on price per generation across image, video, audio and 3D with one request schema for every modality.
How we compare providers
Profiles draw on each provider’s own docs and pricing pages plus third-party reviews and benchmarks, linked as sources on every profile. Subconscious is one of the providers compared, so each head-to-head also says where the other side wins. Pricing and model lineups change often; treat figures as a snapshot from the date above.
By The Subconscious Team · Updated