# Google Vertex AI vs Modal

> Vertex AI hands you hosted models and an MLOps stack. Modal hands you serverless GPUs for Python and leaves the models to you, billed by the second.

Canonical: https://www.subconscious.dev/compare/google-vertex-vs-modal · By The Subconscious Team · Updated September 30, 2026

## How they compare

Vertex AI and Modal sit at different layers. Vertex is a managed model platform: pick Gemini 3.8, Claude or one of 200+ models from Model Garden and call it, or train your own inside Google's pipelines and registry. Modal has no model catalog and no per-token price. A developer decorates a Python function with the GPU it needs, and Modal builds the container, autoscales it and scales it to zero, billing per second from $0.59 an hour for a T4 to $3.95 for an H100 at list. Vertex manages the models. Modal manages the compute.

That makes them complementary more often than competing. A team can call Gemini and Claude through Vertex, then run custom models, embeddings, OCR or transcription on Modal. For custom work, Modal is lighter to adopt and generous to start, with $30 of free credits every month. Costs climb, though: non-preemptible US production runs about 3.75x list, and keeping containers warm to dodge cold starts turns the serverless bill into an always-on one. Vertex's lock-in is heavier, but it adds governance and data integration that Modal does not try to provide.

## What each one does

### Google Vertex AI

Vertex AI is Google Cloud's enterprise AI platform. At Google Cloud Next on April 22, 2026, Google rebranded it the Gemini Enterprise Agent Platform with an agent-first structure, though the API endpoint and most docs still say Vertex. Model Garden offers 200+ models, including Google's Gemini 3.8 family, Anthropic's Claude models and open models like Gemma, alongside Imagen, Veo and Chirp for media and speech. Google's own TPUs sit underneath much of its first-party serving.

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

## Which is best, and when

### Choose Google Vertex AI for

- Hosted Gemini and Claude with no serving code
- Governed training pipelines on Google Cloud
- Enterprise agents with managed memory

### Choose Modal for

- Bursty GPU jobs like embeddings, reranking and transcription
- Private or fine-tuned models on your own serving code
- Python teams shipping a GPU service in an afternoon

## At a glance

| Attribute | Google Vertex AI | Modal |
|---|---|---|
| Model access | Closed and open, 200+ models | Bring your own weights |
| Flagship models | Gemini 3.8 Flash, Claude, Gemma | None hosted |
| Speed | Flash tier built for low latency | ~1s container boot |
| Price | Gemini 3.8 Flash $0.75 in, $3.75 out | Per second; H100 $3.95/hr list |
| Customization | Custom training on GPUs or TPUs | Run any training code |
| Deployment | Managed on Google Cloud | Serverless GPU containers |
| Long context | 1M on Gemini 3.8 Flash | Depends on the model you deploy |

## FAQ

### What is the difference between Google Vertex AI and Modal?

Vertex AI hands you hosted models and an MLOps stack. Modal hands you serverless GPUs for Python and leaves the models to you, billed by the second.

### When should I choose Google Vertex AI over Modal?

Hosted Gemini and Claude with no serving code; Governed training pipelines on Google Cloud; Enterprise agents with managed memory.

### When should I choose Modal over Google Vertex AI?

Bursty GPU jobs like embeddings, reranking and transcription; Private or fine-tuned models on your own serving code; Python teams shipping a GPU service in an afternoon.

### Is Google Vertex AI or Modal cheaper?

Google Vertex AI: Gemini 3.8 Flash $0.75 in, $3.75 out. Modal: Per second; H100 $3.95/hr list. The cheaper choice depends on the model and workload.

### Which has more context, Google Vertex AI or Modal?

Google Vertex AI: 1M on Gemini 3.8 Flash. Modal: Depends on the model you deploy.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Google Vertex AI](https://www.subconscious.dev/compare/subconscious-vs-google-vertex.md), [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md).

Full profiles: [Google Vertex AI](https://www.subconscious.dev/providers/google-vertex.md), [Modal](https://www.subconscious.dev/providers/modal.md).
