Long-running agents deserve better inference.
vs

Fireworks AI vs Luminal

Both sell speed on open models through custom serving. Fireworks is an established host with fine-tuning; Luminal is a young compiler startup.

By The Subconscious Team · Updated

Fireworks AI vs Luminal: key differences

Fireworks runs its own tuned serving stack across 400+ models, with third-party measurements of 167 to 174 tokens per second on DeepSeek V4 Pro, plus SFT, DPO and RFT with fine-tunes served at base price. Luminal reaches for speed through compilation instead: models become a graph of 15 primitive ops, and a search-based compiler emits fused native kernels ahead of time. Luminal reports GPT-OSS 120B at 36K tokens per second on 8 H100s versus 28K for TensorRT-LLM.

Fireworks is the proven choice with a public price sheet, enterprise certifications and managed training. Luminal's cloud is early access with no public catalog or rates, and its benchmark is aggregate throughput rather than per-user speed. Where Luminal stands apart is ownership: the compiler is open source and can be licensed on-prem, which Fireworks' stack is not.

What Fireworks AI and Luminal do

Fireworks AI

Fireworks AI was founded in 2022 by former Meta PyTorch engineers led by CEO Lin Qiao, and it sells speed on open models. Its custom serving stack has posted 167 to 174 tokens per second on DeepSeek V4 Pro in third-party measurements, several times what most GPU peers hit on the same model. The catalog holds 400+ models across text, vision, audio and embeddings, served through an OpenAI-compatible API. In July 2026 it raised a $1.505B Series D at a $17.5B valuation, with a reported $1B+ run rate and 40T+ tokens a day.

Example models: DeepSeek V4 Pro, Kimi K3

Full Fireworks AI profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Fireworks AI or Luminal?

Fireworks AI

Choose Fireworks AI for

  • Independently measured speed on popular open models
  • Managed SFT, DPO and RFT
  • Public pricing and enterprise certifications

Luminal

Choose Luminal for

  • An open-source engine teams can run on their own hardware
  • On-prem deployments with custom kernel work and SLAs
  • Serving custom or fine-tuned architectures off any catalog

Fireworks AI vs Luminal at a glance

AttributeFireworks AILuminal
Model accessOpen weightsBring your own weights
Flagship modelsDeepSeek V4 Pro, Kimi K3No public catalog
Speed167–174 tok/s on DeepSeek V4 Pro36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
PriceFine-tunes served at base pricePay per use; rates not published
CustomizationSFT, DPO, RFT; Training APICompiles any PyTorch or HF model
DeploymentServerless, dedicated GPUsServerless (early access), on-prem license
Long contextFull 1M on DeepSeek V4 ProUnknown

Frequently asked questions

What is the difference between Fireworks AI and Luminal?

Both sell speed on open models through custom serving. Fireworks is an established host with fine-tuning; Luminal is a young compiler startup.

When should I choose Fireworks AI over Luminal?

Independently measured speed on popular open models; Managed SFT, DPO and RFT; Public pricing and enterprise certifications.

When should I choose Luminal over Fireworks AI?

An open-source engine teams can run on their own hardware; On-prem deployments with custom kernel work and SLAs; Serving custom or fine-tuned architectures off any catalog.

Is Fireworks AI or Luminal cheaper?

Fireworks AI: Fine-tunes served at base price. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.