We raised $5.1M for long-running agents.
vs

Modal vs Cloudflare Workers AI

Both are serverless, but Modal rents GPUs by the second for your own code, while Cloudflare Workers AI rents finished models by the token.

By The Subconscious Team · Updated

Modal vs Cloudflare Workers AI: key differences

The unit of purchase decides this. Modal sells compute: decorate a Python function with the GPU it needs, and Modal builds, schedules and autoscales the container, billing per second from $0.59 an hour for a T4 to $3.95 for an H100 at list. It has no model catalog, so teams bring their own weights and serving code. Cloudflare sells inference on 50+ hosted open models, with per-token equivalents like $1.32 in and $3.96 out on DeepSeek V4 Pro and $0.35 in and $0.75 out on gpt-oss 120B. Free allowances differ in shape. Modal's Starter plan renews $30 of credits monthly, while Cloudflare gives 10,000 Neurons a day.

Customization is Modal's clear win. It runs any training or serving code on up to 8 GPUs per container, from T4 through B300, which covers fine-tunes, embeddings, OCR, transcription and batch jobs. Cloudflare limits customization to bring-your-own LoRA on smaller models. Modal's costs rise fast in production, though: non-preemptible US capacity runs about 3.75x list, putting an H100 near $14.81 an hour, and keeping containers warm to avoid weight-loading cold starts turns the bill always-on. Cloudflare needs no ops and no warm pools, though large models can queue. Use Modal for private models and GPU jobs. Use Cloudflare when a standard open model does the job.

What Modal and Cloudflare Workers AI do

Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper

Full Modal profile

Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

Example models: DeepSeek V4 Pro, GLM 5.3

Full Cloudflare Workers AI profile

Should you choose Modal or Cloudflare Workers AI?

Modal

Choose Modal for

  • Private or fine-tuned models on your own serving code
  • Bursty GPU jobs like transcription and batch embeddings
  • Training and inference on one platform

Cloudflare Workers AI

Choose Cloudflare Workers AI for

  • Hosted open LLMs with no containers to manage
  • Per-token billing without warm-pool costs
  • Models called directly from Workers

Modal vs Cloudflare Workers AI at a glance

AttributeModalCloudflare Workers AI
Model accessBring your own weightsOpen weights
Flagship modelsNone hostedDeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B
Speed~1s container bootUnknown
PricePer second; H100 $3.95/hr list$0.011 per 1K Neurons; 10K free daily
CustomizationRun any training codeBYO LoRA on small models (beta)
DeploymentServerless GPU containersServerless on Cloudflare network
Long contextDepends on the model you deploy1M on DeepSeek V4; 262K on Kimi

Frequently asked questions

What is the difference between Modal and Cloudflare Workers AI?

Both are serverless, but Modal rents GPUs by the second for your own code, while Cloudflare Workers AI rents finished models by the token.

When should I choose Modal over Cloudflare Workers AI?

Private or fine-tuned models on your own serving code; Bursty GPU jobs like transcription and batch embeddings; Training and inference on one platform.

When should I choose Cloudflare Workers AI over Modal?

Hosted open LLMs with no containers to manage; Per-token billing without warm-pool costs; Models called directly from Workers.

Is Modal or Cloudflare Workers AI cheaper?

Modal: Per second; H100 $3.95/hr list. Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. The cheaper choice depends on the model and workload.

Which has more context, Modal or Cloudflare Workers AI?

Modal: Depends on the model you deploy. Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.