Long-running agents deserve better inference.
vs

Google Vertex AI vs Luminal

Vertex AI is Google Cloud's full model platform. Luminal is a focused compiler startup that makes open models run faster on GPUs.

By The Subconscious Team · Updated

Google Vertex AI vs Luminal: key differences

Vertex AI bundles Gemini, Claude, Gemma and 200+ other models with training on GPUs or TPUs, pipelines, registries and the rest of Google Cloud. Its strength is breadth inside one enterprise account. Luminal does one thing: it compiles a PyTorch or Hugging Face model into native kernels ahead of time, then serves it serverless in early access or licenses the engine for on-prem use.

Vertex suits teams already on Google Cloud that want managed models plus MLOps. Luminal suits teams that run their own open model and care most about throughput per GPU; it reports GPT-OSS 120B at 36K tokens per second on 8 H100s, about 1.4x vLLM on its own benchmark. Luminal brings no closed models, no training stack and no published prices.

What Google Vertex AI and Luminal do

Google Vertex AI

Vertex AI is Google Cloud's enterprise AI platform. At Google Cloud Next on April 22, 2026, Google rebranded it the Gemini Enterprise Agent Platform with an agent-first structure, though the API endpoint and most docs still say Vertex. Model Garden offers 200+ models, including Google's Gemini 3.8 family, Anthropic's Claude models and open models like Gemma, alongside Imagen, Veo and Chirp for media and speech. Google's own TPUs sit underneath much of its first-party serving.

Example models: Gemini 3.8, Claude

Full Google Vertex AI profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Google Vertex AI or Luminal?

Google Vertex AI

Choose Google Vertex AI for

  • Gemini and Claude under one Google Cloud bill
  • Custom training on TPUs
  • End-to-end MLOps

Luminal

Choose Luminal for

  • Maximum throughput per GPU on a self-chosen model
  • On-prem deployments with custom kernel work and SLAs
  • An open-source engine teams can run on their own hardware

Google Vertex AI vs Luminal at a glance

AttributeGoogle Vertex AILuminal
Model accessClosed and open, 200+ modelsBring your own weights
Flagship modelsGemini 3.8 Flash, Claude, GemmaNo public catalog
SpeedFlash tier built for low latency36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
PriceGemini 3.8 Flash $0.75 in, $3.75 outPay per use; rates not published
CustomizationCustom training on GPUs or TPUsCompiles any PyTorch or HF model
DeploymentManaged on Google CloudServerless (early access), on-prem license
Long context1M on Gemini 3.8 FlashUnknown

Frequently asked questions

What is the difference between Google Vertex AI and Luminal?

Vertex AI is Google Cloud's full model platform. Luminal is a focused compiler startup that makes open models run faster on GPUs.

When should I choose Google Vertex AI over Luminal?

Gemini and Claude under one Google Cloud bill; Custom training on TPUs; End-to-end MLOps.

When should I choose Luminal over Google Vertex AI?

Maximum throughput per GPU on a self-chosen model; On-prem deployments with custom kernel work and SLAs; An open-source engine teams can run on their own hardware.

Is Google Vertex AI or Luminal cheaper?

Google Vertex AI: Gemini 3.8 Flash $0.75 in, $3.75 out. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.