# Baseten vs Cloudflare Workers AI

> Baseten pairs a 13-model API with dedicated GPUs for anything you package. Cloudflare Workers AI offers a wider serverless catalog but no dedicated option for large models.

Canonical: https://www.subconscious.dev/compare/baseten-vs-cloudflare-workers-ai · By The Subconscious Team · Updated September 30, 2026

## How they compare

Baseten runs a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, and posted the lowest time to first token on the Artificial Analysis board in August 2026 at 0.49 seconds. Its endpoints speak both the OpenAI and Anthropic Messages shapes, so Claude Code or an Anthropic SDK can point at it directly. Cloudflare lists 50+ models, including GLM 5.3 and Kimi K2.7 Code, behind OpenAI-compatible endpoints. It publishes per-token prices like $0.35 in and $0.75 out for gpt-oss 120B and gives 10,000 Neurons a day free, but no speed figures. For agentic coding traffic, Baseten's KV cache-aware routing competes with Cloudflare's prefix caching and session-affinity header.

The biggest split is custom models. Baseten's dedicated deployments take anything packaged with its open-source Truss CLI, including speech and embedding models, billed per GPU minute with an H100 near $6.50 an hour, scale to zero and a 99.99% uptime SLA. It also offers self-hosting, HIPAA and data residency, and white-label APIs for model labs. Cloudflare's customization stops at bring-your-own LoRA on smaller models, and large models need the Workers Paid plan. Cloudflare is the cheaper, lower-effort choice for standard open models inside Workers apps. Baseten is the better fit when you need private fine-tunes served, tight latency, or a model outside either catalog.

## What each one does

### Baseten

Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.

### Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

## Which is best, and when

### Choose Baseten for

- Serving private fine-tunes on dedicated GPUs
- Claude Code and Anthropic SDK users on open models
- Lowest measured time to first token

### Choose Cloudflare Workers AI for

- Broader serverless catalog with no GPU management
- Free daily usage for low-traffic features
- Kimi K2.7 Code and GLM 5.3 inside Workers apps

## At a glance

| Attribute | Baseten | Cloudflare Workers AI |
|---|---|---|
| Model access | Open weights, 13 curated | Open weights |
| Flagship models | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B |
| Speed | 0.49s TTFT, lowest measured | - |
| Price | H100 about $6.50/hr dedicated | $0.011 per 1K Neurons; 10K free daily |
| Customization | Deploy any model with Truss | BYO LoRA on small models (beta) |
| Deployment | Model APIs, dedicated, self-host | Serverless on Cloudflare network |
| Long context | Varies by model | 1M on DeepSeek V4; 262K on Kimi |

## FAQ

### What is the difference between Baseten and Cloudflare Workers AI?

Baseten pairs a 13-model API with dedicated GPUs for anything you package. Cloudflare Workers AI offers a wider serverless catalog but no dedicated option for large models.

### When should I choose Baseten over Cloudflare Workers AI?

Serving private fine-tunes on dedicated GPUs; Claude Code and Anthropic SDK users on open models; Lowest measured time to first token.

### When should I choose Cloudflare Workers AI over Baseten?

Broader serverless catalog with no GPU management; Free daily usage for low-traffic features; Kimi K2.7 Code and GLM 5.3 inside Workers apps.

### Is Baseten or Cloudflare Workers AI cheaper?

Baseten: H100 about $6.50/hr dedicated. Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. The cheaper choice depends on the model and workload.

### Which has more context, Baseten or Cloudflare Workers AI?

Baseten: Varies by model. Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Baseten](https://www.subconscious.dev/compare/subconscious-vs-baseten.md), [Subconscious vs Cloudflare Workers AI](https://www.subconscious.dev/compare/subconscious-vs-cloudflare-workers-ai.md).

Full profiles: [Baseten](https://www.subconscious.dev/providers/baseten.md), [Cloudflare Workers AI](https://www.subconscious.dev/providers/cloudflare-workers-ai.md).
