We raised $5.1M for long-running agents.

Mistral AI

European lab shipping open-weight models on its own API, every major cloud, or your GPUs.

Founded
2023
Example models
Mistral Medium 3.5, Mistral Small 4

By The Subconscious Team · Updated

What is Mistral AI?

Mistral AI is a Paris lab that sells its models through La Plateforme, its own API, and releases most of them as open weights. It consolidated the lineup in 2026. Mistral Medium 3.5, released April 28, is a dense 128B model that merges instruction following, reasoning and coding into one set of weights, and it replaced both Devstral 2 and the Magistral reasoning models. It costs $1.50 in and $7.50 out per million tokens and scores 77.6% on SWE-Bench Verified by Mistral's count. Mistral Small 4, a 119B mixture-of-experts model with 6.5B active, costs $0.15 in and $0.60 out. Mistral Large 3, a 675B MoE under Apache 2.0, runs $0.50 in and $1.50 out. All three carry a 256K context window.

Around the general models, Mistral sells Codestral for low-latency fill-in-the-middle completion at $0.30 in and $0.90 out with 128K context, plus OCR, Voxtral speech models and an Agents API with built-in tools. Batch halves prices and cached input cuts input cost by up to 90%. The same models run on Azure AI, Amazon Bedrock, Vertex AI, Snowflake Cortex and IBM watsonx, and Medium 3.5 self-hosts on as few as four GPUs, with NVIDIA NIM containers available. Regional endpoints in Europe and the US went GA in August 2026 alongside a Priority Tier with uptime SLAs. The self-serve fine-tuning API is deprecated; custom training now goes through Forge, an enterprise system covering pre-training, post-training and RL.

Mistral AI pros, cons and use cases

Upsides

  • Open weights on the flagship models, so the same model can move from API to cloud to self-hosted.
  • Low list prices, with Large 3 at $0.50 in and Small 4 at $0.15 in, plus 50% off on Batch.
  • Choice of EU or US processing region, which suits teams with data residency rules.
  • Available on Azure, Bedrock, Vertex AI, Snowflake and watsonx for cloud-credit buyers.

Core use cases

  • Agentic coding and long-horizon tool use on Medium 3.5.
  • Self-hosted or sovereign deployments that need open weights and in-region processing.
  • High-volume, cost-sensitive workloads on Small 4.

Downsides

  • Context tops out at 256K, well below the 1M windows offered by several US labs.
  • Models retire fast: Devstral 2 and Magistral were deprecated within months of launch, which forces regular migrations.

Mistral AI alternatives compared

Pick any row for the full head-to-head.

ProviderModel accessFlagship modelsSpeedPriceCustomizationDeploymentLong contextCompare
Mistral AIOpen weights, plus closed CodestralMistral Medium 3.5, Small 4, Large 3Unknown$0.15–$1.50 in, $0.60–$7.50 out per 1MForge (enterprise); fine-tuning API deprecatedAPI, Azure, Bedrock, Vertex, self-host256K
SubconsciousOpen weightsGLM 5.3, DeepSeek V4.1 Flash2x faster task completion50–80% lower cost; billed on processed tokensMarathon post-trained variantsManaged API, dedicated, on-prem5M+ effective contextCompare
OpenAIClosed, plus open gpt-ossGPT-6 Astra, GPT-5.6 Sol, Terra, LunaFast mode: up to 2.5x at 2x price$0.20–$10 in, $1.20–$50 out per 1MN/AAPI, Azure OpenAI, Bedrock1.05M; 2x input past 272KCompare
AnthropicClosedClaude Fable 5.1, Opus, Sonnet, Haiku 4.5Fable is the slowest tier$1–$10 in, $5–$50 out per 1MN/AAPI, Bedrock, Vertex AI, Microsoft Foundry1M, no surcharge past 200KCompare
Google Vertex AIClosed and open, 200+ modelsGemini 3.8 Flash, Claude, GemmaFlash tier built for low latencyGemini 3.8 Flash $0.75 in, $3.75 outCustom training on GPUs or TPUsManaged on Google Cloud1M on Gemini 3.8 FlashCompare
Amazon BedrockClosed and open, 100+ modelsClaude, GPT-6 Astra, Nova, DeepSeekLatency-optimized option on some models~20–35% above direct; Claude at parityFine-tuning, Custom Model ImportManaged on AWS, AgentCoreVaries by modelCompare
Together AIOpen weightsKimi K3, DeepSeek V4, GLM 5.2, Qwen 3.80.99s TTFT on DeepSeek V4 ProParity with Fireworks and BasetenLoRA and full SFT; RL in betaServerless, dedicated, GPU clusters512K on DeepSeek V4 ProCompare
Fireworks AIOpen weightsDeepSeek V4 Pro, Kimi K3167–174 tok/s on DeepSeek V4 ProFine-tunes served at base priceSFT, DPO, RFT; Training APIServerless, dedicated GPUsFull 1M on DeepSeek V4 ProCompare
BasetenOpen weights, 13 curatedGLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B0.49s TTFT, lowest measuredH100 about $6.50/hr dedicatedDeploy any model with TrussModel APIs, dedicated, self-hostVaries by modelCompare
GroqOpen weightsGPT-OSS 120B, Qwen 3.6 27B500–1,000 tok/sNear the floor on small modelsNo fine-tuned model hostingGroqCloud APIAround 131K maxCompare
CerebrasOpen weightsGPT-OSS 120B, Gemma 4 31B~3,000 tok/s on GPT-OSS 120B$0.35 in, $0.75 out (GPT-OSS 120B)UnknownShared API, dedicated, partnersUnknownCompare
DeepInfraOpen weightsDeepSeek V4 Flash, Llama 3.1 8B~33 tok/s on DeepSeek V4 Pro (FP4)From $0.02 per 1MNo managed fine-tuningShared API, no contracts66K on FP4 DeepSeek V4 ProCompare
Hugging Face Inference ProvidersOpen weightsGLM 5.3, Kimi K3, DeepSeek V4.1 FlashRoutes to fastest provider by defaultProvider rates, no markupN/AServerless router; dedicated EndpointsUp to 1M, provider-dependentCompare
ModalBring your own weightsNone hosted~1s container bootPer second; H100 $3.95/hr listRun any training codeServerless GPU containersDepends on the model you deployCompare
Cloudflare Workers AIOpen weightsDeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120BUnknown$0.011 per 1K Neurons; 10K free dailyBYO LoRA on small models (beta)Serverless on Cloudflare network1M on DeepSeek V4; 262K on KimiCompare
xAIClosedGrok 4.6, Grok 4.20, grok-build~54 tok/s on Grok 4.6$2 in, $6 out (Grok 4.6); 2x past 200KUnknownFirst-party API500K (4.6), 1M (4.20, 4.3)Compare
DeepSeekOpen weights (MIT)DeepSeek V4.1 Flash, V4 Pro~35 tok/s on V4 ProOff-peak hours at half priceOpen weights to fine-tuneFirst-party API, Hugging Face weights1M, 384K max outputCompare
Moonshot AIOpen weights, custom licenseKimi K3, Kimi K2.6~33 tok/s on Kimi K3$3 in, $15 out (Kimi K3)Open weights to fine-tuneAPI, Kimi Code, OpenRouter1MCompare
Z.aiOpen weights (MIT)GLM-5.3, GLM-5.3-Flash~80 tok/s on GLM-5.3$1.40 in, $4.40 out (GLM-5.3); free Flash tierOpen weights, no license limitsAPI, GLM Coding Plan1M (GLM-5.3)Compare
Alibaba CloudClosed Max; open smaller QwenQwen 3.8-Max, Qwen 3.7-Max~40 tok/s on Qwen 3.8-Max$2 in, $6 out internationalNo fine-tuning on MaxModel Studio on Alibaba Cloud1M (Qwen 3.8-Max)Compare
MetaClosed API; open Muse GlimmerMuse Spark 1.3, Muse Glimmer~145–233 tok/s on Muse Spark 1.3$1.25 in, $4.25 out; Contributor tier cheaperOpen Muse Glimmer weights to fine-tuneMeta Model API (preview)1MCompare
CohereClosed, plus open Command A+Command A+, Command A, Embed 4, Rerank 4375 tok/s on Command A+ W4A4, per Cohere$0.0375–$2.50 in, $0.15–$10 out per 1MEnterprise fine-tuning, incl. privateAPI, Bedrock, Azure, OCI, VPC, on-prem256K on Command A; 128K on A+Compare
SambaNovaOpen weightsMiniMax M2.7, GPT-OSS 120B, DeepSeek~820 tok/s on MiniMax M2.7 (SN50)$0.22 in, $0.59 out (GPT-OSS 120B)UnknownSambaCloud, racks for neocloudsUp to 192K (MiniMax M2.7)Compare
NebiusOpen weights, 60+ modelsDeepSeek, Qwen, GLM, Kimi, GPT-OSSAmong top hosts on throughputFrom $0.06 per 1M inputServe uploaded fine-tunesToken Factory, dedicated, raw GPUsVaries by modelCompare
CrusoeOpen weightsDeepSeek V4, GLM 5.3, Kimi K2.6, Nemotron 3Up to 9.9x faster TTFT vs vLLM (vendor claim)$0.05–$1.74 in, $0.20–$4.40 out per 1MServerless LoRA fine-tuningServerless, self-serve and tailored dedicated, raw GPUsVaries by model; cluster-wide KV cacheCompare
falHosted media modelsFLUX, Kling, SeedreamCold starts on less popular endpointsPer image, per video second, GPU timeLoRA training endpointsHosted API, serverless GPUsNot applicableCompare
Novita AIOpen weightsDeepSeek V4 Pro, Gemma 4~36 tok/s on DeepSeek V4 ProFrom $0.02 per 1M; batch 50% offHot-swappable LoRA adaptersServerless, GPU cloud, dedicatedFull 1M on DeepSeek V4 ProCompare
VeniceOpen weights, plus proxied closed modelsGLM 5.3, Kimi K3, DeepSeek V4 ProUnknown$0.06–$12 in, $0.28–$60 out per 1M; DIEM stakingUnknownServerless API, consumer app1M on most current modelsCompare
ParasailAny Hugging Face modelGTE-Qwen2, Qwen3-VL-8B-Instruct600ms p99 real-time budgetPer-parameter rates; batch 50% offPrivate Hugging Face reposServerless, elastic, dedicated, batchVaries by modelCompare
Inference.netOpen, closed and customCustomer fine-tunesBatch windows of 24h to 7 daysDiscounted spare GPU capacityDistill traces into custom modelsBatch API, gateway, dedicated GPUsVaries by modelCompare
GMI CloudOpen and third-party modelsGLM-4.7-Flash, Google VeoNear bare-metal performance$0.07 in, $0.40 out (GLM-4.7-Flash)UnknownShared, autoscaling, reserved GPUsVaries by modelCompare
Thinking MachinesOpen weightsInkling, Inkling-SmallUnknownPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 outLoRA SFT and RL via TinkerTraining API, beta serverless (Inkling only)Inkling up to 1M; Tinker 32K–256KCompare
Sail ResearchOpen weightsKimi K2.6, GLM-5, GPT-OSS 120BMinutes per turn by design30–80% off by completion windowCustomer LoRA fine-tunesAPI plus SailboxesVaries by modelCompare
MorphSpecialist modelsmorph-v3-fast, morph-v3-large10,500+ tok/s Fast Apply~40% fewer tokens than full rewritesFine-tuning offeredOpenAI-compatible APIUnknownCompare
RelaceSpecialist modelsrelace-apply-3, agentic search~10,000 tok/s apply3x+ cheaper than full rewritesUnknownHosted API or self-hosted128K maxCompare
TypeSafe AIDecision modelsJev, jev-1.13~100ms per callA fraction of an LLM callUnknownEarly-access APIUnknownCompare
StepFunOpen (Apache 2.0) and API modelsStep 3.7 Flash, Step3~128 tok/s on Step 3.7 Flash$0.20 in, $1.15 out (Step 3.7 Flash)Open weights to fine-tuneFirst-party API, OpenRouter256KCompare
RunwareHosted media modelsSeedance 2.5, Qwen-Image-3.0UnknownImages from fractions of a centFine-tuned diffusion checkpointsUnified API, raw GPUsNot applicableCompare
StreamLakeProprietary coding modelsKAT-Coder-Pro V2.5, KAT-Coder-AirUnknownPer token or KwaiKAT Coding PlanUnknownMaaS API, bare metalUnknownCompare
WaferOpen weightsQwen 3.5 397B Turbo, GLM 5.1 Turbo2–2.8x vs stock vLLM or SGLangWafer Pass from $10 a weekAgent-tuned dedicated deploymentsServerless pass, dedicatedVaries by modelCompare
RunInfraOpen weightsNemotron 3.5 Lightning 30B, Qwen 3.8 27BCold starts under 2sCoding plans from $10 a monthUploads up to 50 GB; auto-quantizationModel APIs, agent-built endpointsVaries by modelCompare
Particle.AIOpen weightsDeepSeek V4.1 Flash, GLM 5.3 Flash~157 tok/s on DeepSeek V4.1 Flash$0.10 in, $0.40 out (GLM 5.3 Flash)UnknownVia Vercel AI Gateway1MCompare

Frequently asked questions

What is Mistral AI?

Mistral AI is a Paris lab that sells its models through La Plateforme, its own API, and releases most of them as open weights. It consolidated the lineup in 2026. Mistral Medium 3.5, released April 28, is a dense 128B model that merges instruction following, reasoning and coding into one set of weights, and it replaced both Devstral 2 and the Magistral reasoning models. It costs $1.50 in and $7.50 out per million tokens and scores 77.6% on SWE-Bench Verified by Mistral's count. Mistral Small 4, a 119B mixture-of-experts model with 6.5B active, costs $0.15 in and $0.60 out. Mistral Large 3, a 675B MoE under Apache 2.0, runs $0.50 in and $1.50 out. All three carry a 256K context window.

What is Mistral AI best for?

Agentic coding and long-horizon tool use on Medium 3.5; Self-hosted or sovereign deployments that need open weights and in-region processing; High-volume, cost-sensitive workloads on Small 4.

How much does Mistral AI cost?

Mistral AI pricing at a glance: $0.15–$1.50 in, $0.60–$7.50 out per 1M. Rates change often, so check Mistral AI's pricing page before committing.

How much context does Mistral AI support?

Mistral AI's long-context support: 256K.

What are the downsides of Mistral AI?

Context tops out at 256K, well below the 1M windows offered by several US labs; Models retire fast: Devstral 2 and Magistral were deprecated within months of launch, which forces regular migrations.

What are the best alternatives to Mistral AI?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Mistral AI on this site.

Sources: Mistral pricing, Mistral Medium 3.5 launch, Mistral models overview, Mistral regional inference. Pricing and model lineups change often; figures are a snapshot.