Long-running agents deserve better inference.
vs

Mistral AI vs Luminal

Mistral ships open-weight models on its API, clouds or your GPUs. Luminal makes models like Mistral run faster on that hardware.

By The Subconscious Team · Updated

Mistral AI vs Luminal: key differences

Mistral is a European lab with open-weight models such as Mistral Medium 3.5, Small 4 and Large 3, available on its API, Azure, Bedrock, Vertex or self-hosted, with 256K context. Luminal does not train models. It compiles a model into native kernels ahead of time and sells that as serverless endpoints in early access or an on-prem license.

They can stack: a team self-hosting Mistral weights could serve them through Luminal's compiler instead of vLLM. Luminal reports GPT-OSS 120B at 36K tokens per second on 8 H100s against 26K for vLLM, measured in-house. Choose Mistral for the models and European provenance, and consider Luminal for how they run.

What Mistral AI and Luminal do

Mistral AI

Mistral AI is a Paris lab that sells its models through La Plateforme, its own API, and releases most of them as open weights. It consolidated the lineup in 2026. Mistral Medium 3.5, released April 28, is a dense 128B model that merges instruction following, reasoning and coding into one set of weights, and it replaced both Devstral 2 and the Magistral reasoning models. It costs $1.50 in and $7.50 out per million tokens and scores 77.6% on SWE-Bench Verified by Mistral's count. Mistral Small 4, a 119B mixture-of-experts model with 6.5B active, costs $0.15 in and $0.60 out. Mistral Large 3, a 675B MoE under Apache 2.0, runs $0.50 in and $1.50 out. All three carry a 256K context window.

Example models: Mistral Medium 3.5, Mistral Small 4

Full Mistral AI profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Mistral AI or Luminal?

Mistral AI

Choose Mistral AI for

  • European open-weight models
  • Deployment across every major cloud
  • Enterprise customization through Forge

Luminal

Choose Luminal for

  • Serving self-hosted open weights faster than vLLM
  • On-prem deployments with custom kernel work and SLAs
  • An open-source engine teams can run on their own hardware

Mistral AI vs Luminal at a glance

AttributeMistral AILuminal
Model accessOpen weights, plus closed CodestralBring your own weights
Flagship modelsMistral Medium 3.5, Small 4, Large 3No public catalog
SpeedUnknown36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price$0.15–$1.50 in, $0.60–$7.50 out per 1MPay per use; rates not published
CustomizationForge (enterprise); fine-tuning API deprecatedCompiles any PyTorch or HF model
DeploymentAPI, Azure, Bedrock, Vertex, self-hostServerless (early access), on-prem license
Long context256KUnknown

Frequently asked questions

What is the difference between Mistral AI and Luminal?

Mistral ships open-weight models on its API, clouds or your GPUs. Luminal makes models like Mistral run faster on that hardware.

When should I choose Mistral AI over Luminal?

European open-weight models; Deployment across every major cloud; Enterprise customization through Forge.

When should I choose Luminal over Mistral AI?

Serving self-hosted open weights faster than vLLM; On-prem deployments with custom kernel work and SLAs; An open-source engine teams can run on their own hardware.

Is Mistral AI or Luminal cheaper?

Mistral AI: $0.15–$1.50 in, $0.60–$7.50 out per 1M. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.