vs

Google Vertex AI vs Modal

Vertex AI hands you hosted models and an MLOps stack. Modal hands you serverless GPUs for Python and leaves the models to you, billed by the second.

By The Subconscious Team · Updated

Google Vertex AI vs Modal: key differences

Vertex AI and Modal sit at different layers. Vertex is a managed model platform: pick Gemini 3.8, Claude or one of 200+ models from Model Garden and call it, or train your own inside Google's pipelines and registry. Modal has no model catalog and no per-token price. A developer decorates a Python function with the GPU it needs, and Modal builds the container, autoscales it and scales it to zero, billing per second from $0.59 an hour for a T4 to $3.95 for an H100 at list. Vertex manages the models. Modal manages the compute.

That makes them complementary more often than competing. A team can call Gemini and Claude through Vertex, then run custom models, embeddings, OCR or transcription on Modal. For custom work, Modal is lighter to adopt and generous to start, with $30 of free credits every month. Costs climb, though: non-preemptible US production runs about 3.75x list, and keeping containers warm to dodge cold starts turns the serverless bill into an always-on one. Vertex's lock-in is heavier, but it adds governance and data integration that Modal does not try to provide.

What Google Vertex AI and Modal do

Google Vertex AI

Vertex AI is Google Cloud's enterprise AI platform. At Google Cloud Next on April 22, 2026, Google rebranded it the Gemini Enterprise Agent Platform with an agent-first structure, though the API endpoint and most docs still say Vertex. Model Garden offers 200+ models, including Google's Gemini 3.8 family, Anthropic's Claude models and open models like Gemma, alongside Imagen, Veo and Chirp for media and speech. Google's own TPUs sit underneath much of its first-party serving.

Example models: Gemini 3.8, Claude

Full Google Vertex AI profile

Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper

Full Modal profile

Should you choose Google Vertex AI or Modal?

Google Vertex AI

Choose Google Vertex AI for

  • Hosted Gemini and Claude with no serving code
  • Governed training pipelines on Google Cloud
  • Enterprise agents with managed memory

Modal

Choose Modal for

  • Bursty GPU jobs like embeddings, reranking and transcription
  • Private or fine-tuned models on your own serving code
  • Python teams shipping a GPU service in an afternoon

Google Vertex AI vs Modal at a glance

AttributeGoogle Vertex AIModal
Model accessClosed and open, 200+ modelsBring your own weights
Flagship modelsGemini 3.8 Flash, Claude, GemmaNone hosted
SpeedFlash tier built for low latency~1s container boot
PriceGemini 3.8 Flash $0.75 in, $3.75 outPer second; H100 $3.95/hr list
CustomizationCustom training on GPUs or TPUsRun any training code
DeploymentManaged on Google CloudServerless GPU containers
Long context1M on Gemini 3.8 FlashDepends on the model you deploy

Frequently asked questions

What is the difference between Google Vertex AI and Modal?

Vertex AI hands you hosted models and an MLOps stack. Modal hands you serverless GPUs for Python and leaves the models to you, billed by the second.

When should I choose Google Vertex AI over Modal?

Hosted Gemini and Claude with no serving code; Governed training pipelines on Google Cloud; Enterprise agents with managed memory.

When should I choose Modal over Google Vertex AI?

Bursty GPU jobs like embeddings, reranking and transcription; Private or fine-tuned models on your own serving code; Python teams shipping a GPU service in an afternoon.

Is Google Vertex AI or Modal cheaper?

Google Vertex AI: Gemini 3.8 Flash $0.75 in, $3.75 out. Modal: Per second; H100 $3.95/hr list. The cheaper choice depends on the model and workload.

Which has more context, Google Vertex AI or Modal?

Google Vertex AI: 1M on Gemini 3.8 Flash. Modal: Depends on the model you deploy.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.