# Subconscious vs Cloudflare Workers AI

> Both serve GLM 5.3, but for different jobs. Cloudflare Workers AI puts 50+ open models next to Workers code; Subconscious is built for agent traces that run past 200K tokens.

Canonical: https://www.subconscious.dev/compare/subconscious-vs-cloudflare-workers-ai · By The Subconscious Team · Updated September 30, 2026

## How they compare

The overlap is concrete. Both serve GLM 5.3, which Cloudflare lists at $1.40 in and $4.40 out per million tokens. Cloudflare also carries DeepSeek V4 Pro with the full 1,048,576 token context, and it softens long loops with prefix caching plus an x-session-affinity header that pins a conversation to one instance. Subconscious attacks the same problem in the runtime. It prunes the KV cache instead of rereading a growing context, and against open models on standard inference it delivers 2x faster task completion, a 5M+ effective context and 50 to 80% lower cost. It bills tokens processed after compression, so a long coding trace can bill far fewer tokens than it sends, while Cloudflare bills every input token, discounted only where the prefix cache hits.

Cloudflare wins on breadth and platform. Its catalog lists 50+ open models, including Kimi K2.7 Code, embeddings and gpt-oss 120B at $0.35 in and $0.75 out, and 10,000 Neurons a day are free. Inference, code, storage and the Agents SDK share one platform, and AI Gateway adds caching, retries, fallbacks and spend logs. Subconscious keeps a focused managed catalog of GLM 5.3 and DeepSeek V4.1 Flash, but its dedicated and on-prem deployments run nearly any open model, which Cloudflare does not offer for large LLMs. Short, single-turn requests see little of the Subconscious advantage. The practical split is Cloudflare for chat features and classification inside Workers apps, and Subconscious for long-running coding and research agents.

## What each one does

### Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

## Which is best, and when

### Choose Subconscious for

- Coding agents whose traces run past 200K tokens
- Long GLM 5.3 runs billed on processed tokens
- Dedicated or on-prem serving of other open models

### Choose Cloudflare Workers AI for

- LLM calls made directly from a Worker
- Short classification and embedding jobs within the free daily Neurons
- Picking from 50+ open models on one bill

## At a glance

| Attribute | Subconscious | Cloudflare Workers AI |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B |
| Speed | 2x faster task completion | - |
| Price | 50–80% lower cost; billed on processed tokens | $0.011 per 1K Neurons; 10K free daily |
| Customization | Marathon post-trained variants | BYO LoRA on small models (beta) |
| Deployment | Managed API, dedicated, on-prem | Serverless on Cloudflare network |
| Long context | 5M+ effective context | 1M on DeepSeek V4; 262K on Kimi |

## FAQ

### What is the difference between Subconscious and Cloudflare Workers AI?

Both serve GLM 5.3, but for different jobs. Cloudflare Workers AI puts 50+ open models next to Workers code; Subconscious is built for agent traces that run past 200K tokens.

### When should I choose Subconscious over Cloudflare Workers AI?

Coding agents whose traces run past 200K tokens; Long GLM 5.3 runs billed on processed tokens; Dedicated or on-prem serving of other open models.

### When should I choose Cloudflare Workers AI over Subconscious?

LLM calls made directly from a Worker; Short classification and embedding jobs within the free daily Neurons; Picking from 50+ open models on one bill.

### Is Subconscious or Cloudflare Workers AI cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. The cheaper choice depends on the model and workload.

### Which has more context, Subconscious or Cloudflare Workers AI?

Subconscious: 5M+ effective context. Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi.

## Try Subconscious

Subconscious speaks the OpenAI and Anthropic API formats. Base URL: https://api.subconscious.dev/v1. Docs: https://docs.subconscious.dev. Get an API key: https://platform.subconscious.dev/signin. Agent guide: https://www.subconscious.dev/agents.md.

Full profiles: [Subconscious](https://www.subconscious.dev/providers/subconscious.md), [Cloudflare Workers AI](https://www.subconscious.dev/providers/cloudflare-workers-ai.md).
