# Cloudflare Workers AI vs Thinking Machines

> Thinking Machines sells Tinker, an API for writing your own post-training loops, plus its open Inkling models. Workers AI is a production inference host.

Canonical: https://www.subconscious.dev/compare/cloudflare-workers-ai-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

These solve different stages. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write custom SFT or RL loops on models like Kimi K2.6, GLM-5.3 and gpt-oss while Thinking Machines runs the GPUs. Training is LoRA only. Workers AI's LoRA support is a free beta on small non-quantized models, capped at rank 32 and 300MB, so it is no substitute for serious post-training. On the other side, Tinker's checkpoint sampling endpoint is scoped to testing and low internal traffic, not user-facing load.

For serving, Workers AI covers 50+ models with OpenAI-compatible endpoints, prefix caching, AI Gateway and a free daily tier, and handles 1M context on DeepSeek V4. Thinking Machines' beta serverless API serves only Inkling and Inkling-Small, its Apache 2.0 models with image and audio input and up to 1M context, with Inkling at $1.00 in and $4.05 out. That makes Inkling a notable model, but not a general host. A plausible split: train on Tinker, serve somewhere with dedicated capacity, since Workers AI cannot host large custom weights.

## What each one does

### Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Cloudflare Workers AI for

- Serving production traffic on open LLMs
- Apps that call models from Workers
- Broad catalog with no training work

### Choose Thinking Machines for

- Custom SFT or RL loops on large MoE models
- Fine-tuning Kimi K2.6 or Inkling without a cluster
- Evaluating Inkling with image and audio input

## At a glance

| Attribute | Cloudflare Workers AI | Thinking Machines |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B | Inkling, Inkling-Small |
| Speed | - | - |
| Price | $0.011 per 1K Neurons; 10K free daily | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | BYO LoRA on small models (beta) | LoRA SFT and RL via Tinker |
| Deployment | Serverless on Cloudflare network | Training API, beta serverless (Inkling only) |
| Long context | 1M on DeepSeek V4; 262K on Kimi | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Cloudflare Workers AI and Thinking Machines?

Thinking Machines sells Tinker, an API for writing your own post-training loops, plus its open Inkling models. Workers AI is a production inference host.

### When should I choose Cloudflare Workers AI over Thinking Machines?

Serving production traffic on open LLMs; Apps that call models from Workers; Broad catalog with no training work.

### When should I choose Thinking Machines over Cloudflare Workers AI?

Custom SFT or RL loops on large MoE models; Fine-tuning Kimi K2.6 or Inkling without a cluster; Evaluating Inkling with image and audio input.

### Is Cloudflare Workers AI or Thinking Machines cheaper?

Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Cloudflare Workers AI or Thinking Machines?

Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Cloudflare Workers AI](https://www.subconscious.dev/compare/subconscious-vs-cloudflare-workers-ai.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Cloudflare Workers AI](https://www.subconscious.dev/providers/cloudflare-workers-ai.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
