We raised $5.1M for long-running agents.
vs

Cloudflare Workers AI vs Thinking Machines

Thinking Machines sells Tinker, an API for writing your own post-training loops, plus its open Inkling models. Workers AI is a production inference host.

By The Subconscious Team · Updated

Cloudflare Workers AI vs Thinking Machines: key differences

These solve different stages. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write custom SFT or RL loops on models like Kimi K2.6, GLM-5.3 and gpt-oss while Thinking Machines runs the GPUs. Training is LoRA only. Workers AI's LoRA support is a free beta on small non-quantized models, capped at rank 32 and 300MB, so it is no substitute for serious post-training. On the other side, Tinker's checkpoint sampling endpoint is scoped to testing and low internal traffic, not user-facing load.

For serving, Workers AI covers 50+ models with OpenAI-compatible endpoints, prefix caching, AI Gateway and a free daily tier, and handles 1M context on DeepSeek V4. Thinking Machines' beta serverless API serves only Inkling and Inkling-Small, its Apache 2.0 models with image and audio input and up to 1M context, with Inkling at $1.00 in and $4.05 out. That makes Inkling a notable model, but not a general host. A plausible split: train on Tinker, serve somewhere with dedicated capacity, since Workers AI cannot host large custom weights.

What Cloudflare Workers AI and Thinking Machines do

Cloudflare Workers AI

Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.

Example models: DeepSeek V4 Pro, GLM 5.3

Full Cloudflare Workers AI profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Cloudflare Workers AI or Thinking Machines?

Cloudflare Workers AI

Choose Cloudflare Workers AI for

  • Serving production traffic on open LLMs
  • Apps that call models from Workers
  • Broad catalog with no training work

Thinking Machines

Choose Thinking Machines for

  • Custom SFT or RL loops on large MoE models
  • Fine-tuning Kimi K2.6 or Inkling without a cluster
  • Evaluating Inkling with image and audio input

Cloudflare Workers AI vs Thinking Machines at a glance

AttributeCloudflare Workers AIThinking Machines
Model accessOpen weightsOpen weights
Flagship modelsDeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120BInkling, Inkling-Small
SpeedUnknownUnknown
Price$0.011 per 1K Neurons; 10K free dailyPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationBYO LoRA on small models (beta)LoRA SFT and RL via Tinker
DeploymentServerless on Cloudflare networkTraining API, beta serverless (Inkling only)
Long context1M on DeepSeek V4; 262K on KimiInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Cloudflare Workers AI and Thinking Machines?

Thinking Machines sells Tinker, an API for writing your own post-training loops, plus its open Inkling models. Workers AI is a production inference host.

When should I choose Cloudflare Workers AI over Thinking Machines?

Serving production traffic on open LLMs; Apps that call models from Workers; Broad catalog with no training work.

When should I choose Thinking Machines over Cloudflare Workers AI?

Custom SFT or RL loops on large MoE models; Fine-tuning Kimi K2.6 or Inkling without a cluster; Evaluating Inkling with image and audio input.

Is Cloudflare Workers AI or Thinking Machines cheaper?

Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Cloudflare Workers AI or Thinking Machines?

Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.