# Cloudflare Workers AI vs RunInfra

> RunInfra pairs a small hosted library with an agent that benchmarks and builds custom endpoints. Workers AI offers a ready catalog of frontier open models.

Canonical: https://www.subconscious.dev/compare/cloudflare-workers-ai-vs-runinfra · By The Subconscious Team · Updated September 30, 2026

## How they compare

Both host Qwen 3.8 27B, but beyond that the catalogs diverge. RunInfra's hosted library is tiny and centered on mid-size models such as Nemotron 3.5 Lightning 30B and Ornith 1.5 35B. Workers AI reaches frontier scale with DeepSeek V4 Pro, GLM 5.3 and Kimi K2.7 Code, and 1M context on DeepSeek V4. RunInfra's coding plans start at $10 a month with limits that reset every five hours and weekly, and work with Claude Code, Codex, OpenCode and Aider. Workers AI bills per Neuron with a free daily allocation instead.

Custom deployment is where RunInfra leads. Describe an endpoint in plain English and its agent picks a model, benchmarks it on GPUs from L4 to B200, searches AWQ, GPTQ and FP8 variants and ships an OpenAI-compatible endpoint that scales to zero with cold starts under two seconds. Paid plans take uploads up to 50 GB. Workers AI has no dedicated deployments for large models and limits LoRA to small ones. RunInfra is a 2026 company with little independent benchmarking, while Cloudflare adds AI Gateway, storage and the Agents SDK.

## What each one does

### Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

### RunInfra

RunInfra pitches open models built for agents, with two ways in. Its hosted Model APIs serve a small curated library, including Nemotron 3.5 Lightning 30B, Qwen 3.8 27B and Ornith 1.5 35B, behind one key that works with both the OpenAI and Anthropic SDKs. Cached context bills at a discount. Coding plans start at $10 a month with limits that reset every five hours and every week, and they plug into Claude Code, Codex, OpenCode, Cline, Aider and dozens of other agent CLIs.

## Which is best, and when

### Choose Cloudflare Workers AI for

- Frontier-scale open LLMs without setup
- Apps and agents built on Workers
- Gateway caching and fallbacks

### Choose RunInfra for

- Cheap flat-rate plans for coding CLIs
- Auto-tuned endpoints for uploaded models
- Voice pipelines chaining Whisper, LLM and TTS

## At a glance

| Attribute | Cloudflare Workers AI | RunInfra |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B | Nemotron 3.5 Lightning 30B, Qwen 3.8 27B |
| Speed | - | Cold starts under 2s |
| Price | $0.011 per 1K Neurons; 10K free daily | Coding plans from $10 a month |
| Customization | BYO LoRA on small models (beta) | Uploads up to 50 GB; auto-quantization |
| Deployment | Serverless on Cloudflare network | Model APIs, agent-built endpoints |
| Long context | 1M on DeepSeek V4; 262K on Kimi | Varies by model |

## FAQ

### What is the difference between Cloudflare Workers AI and RunInfra?

RunInfra pairs a small hosted library with an agent that benchmarks and builds custom endpoints. Workers AI offers a ready catalog of frontier open models.

### When should I choose Cloudflare Workers AI over RunInfra?

Frontier-scale open LLMs without setup; Apps and agents built on Workers; Gateway caching and fallbacks.

### When should I choose RunInfra over Cloudflare Workers AI?

Cheap flat-rate plans for coding CLIs; Auto-tuned endpoints for uploaded models; Voice pipelines chaining Whisper, LLM and TTS.

### Is Cloudflare Workers AI or RunInfra cheaper?

Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. RunInfra: Coding plans from $10 a month. The cheaper choice depends on the model and workload.

### Which has more context, Cloudflare Workers AI or RunInfra?

Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi. RunInfra: Varies by model.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Cloudflare Workers AI](https://www.subconscious.dev/compare/subconscious-vs-cloudflare-workers-ai.md), [Subconscious vs RunInfra](https://www.subconscious.dev/compare/subconscious-vs-runinfra.md).

Full profiles: [Cloudflare Workers AI](https://www.subconscious.dev/providers/cloudflare-workers-ai.md), [RunInfra](https://www.subconscious.dev/providers/runinfra.md).
