# Modal vs Cloudflare Workers AI

> Both are serverless, but Modal rents GPUs by the second for your own code, while Cloudflare Workers AI rents finished models by the token.

Canonical: https://www.subconscious.dev/compare/modal-vs-cloudflare-workers-ai · By The Subconscious Team · Updated September 30, 2026

## How they compare

The unit of purchase decides this. Modal sells compute: decorate a Python function with the GPU it needs, and Modal builds, schedules and autoscales the container, billing per second from $0.59 an hour for a T4 to $3.95 for an H100 at list. It has no model catalog, so teams bring their own weights and serving code. Cloudflare sells inference on 50+ hosted open models, with per-token equivalents like $1.32 in and $3.96 out on DeepSeek V4 Pro and $0.35 in and $0.75 out on gpt-oss 120B. Free allowances differ in shape. Modal's Starter plan renews $30 of credits monthly, while Cloudflare gives 10,000 Neurons a day.

Customization is Modal's clear win. It runs any training or serving code on up to 8 GPUs per container, from T4 through B300, which covers fine-tunes, embeddings, OCR, transcription and batch jobs. Cloudflare limits customization to bring-your-own LoRA on smaller models. Modal's costs rise fast in production, though: non-preemptible US capacity runs about 3.75x list, putting an H100 near $14.81 an hour, and keeping containers warm to avoid weight-loading cold starts turns the bill always-on. Cloudflare needs no ops and no warm pools, though large models can queue. Use Modal for private models and GPU jobs. Use Cloudflare when a standard open model does the job.

## What each one does

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

### Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

## Which is best, and when

### Choose Modal for

- Private or fine-tuned models on your own serving code
- Bursty GPU jobs like transcription and batch embeddings
- Training and inference on one platform

### Choose Cloudflare Workers AI for

- Hosted open LLMs with no containers to manage
- Per-token billing without warm-pool costs
- Models called directly from Workers

## At a glance

| Attribute | Modal | Cloudflare Workers AI |
|---|---|---|
| Model access | Bring your own weights | Open weights |
| Flagship models | None hosted | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B |
| Speed | ~1s container boot | - |
| Price | Per second; H100 $3.95/hr list | $0.011 per 1K Neurons; 10K free daily |
| Customization | Run any training code | BYO LoRA on small models (beta) |
| Deployment | Serverless GPU containers | Serverless on Cloudflare network |
| Long context | Depends on the model you deploy | 1M on DeepSeek V4; 262K on Kimi |

## FAQ

### What is the difference between Modal and Cloudflare Workers AI?

Both are serverless, but Modal rents GPUs by the second for your own code, while Cloudflare Workers AI rents finished models by the token.

### When should I choose Modal over Cloudflare Workers AI?

Private or fine-tuned models on your own serving code; Bursty GPU jobs like transcription and batch embeddings; Training and inference on one platform.

### When should I choose Cloudflare Workers AI over Modal?

Hosted open LLMs with no containers to manage; Per-token billing without warm-pool costs; Models called directly from Workers.

### Is Modal or Cloudflare Workers AI cheaper?

Modal: Per second; H100 $3.95/hr list. Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. The cheaper choice depends on the model and workload.

### Which has more context, Modal or Cloudflare Workers AI?

Modal: Depends on the model you deploy. Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md), [Subconscious vs Cloudflare Workers AI](https://www.subconscious.dev/compare/subconscious-vs-cloudflare-workers-ai.md).

Full profiles: [Modal](https://www.subconscious.dev/providers/modal.md), [Cloudflare Workers AI](https://www.subconscious.dev/providers/cloudflare-workers-ai.md).
