Cloudflare Workers AI vs Cohere
Cohere builds for private enterprise RAG, with Embed, Rerank and on-prem deployment. Cloudflare Workers AI serves open models on shared edge infrastructure with public prices.
By The Subconscious Team · Updated
Cloudflare Workers AI vs Cohere: key differences
These two target different buyers. Cohere sells to banks, governments and large enterprises. Command A+, a 218B mixture-of-experts model under Apache 2.0 with 128K context, handles reasoning, vision, translation and tool use, and Command A has a 256K window at $2.50 in and $10 out per million tokens. Command A+ per-token prices are not published, so production often starts with sales. Cloudflare publishes every price: DeepSeek V4 Pro with a full 1M context at $1.32 in and $3.96 out, GLM 5.3 and Kimi K2.7 Code, with 10,000 Neurons a day free. On agentic coding and broad intelligence indexes, Command A+ trails the latest DeepSeek and GLM models that Cloudflare hosts.
Retrieval and private deployment are Cohere's ground. Embed 4 handles text, images and PDFs with a 128K context, and Rerank 4 comes in Pro and Fast versions priced per search. Cohere runs in any VPC or fully on-prem, including fine-tuning inside that environment, and its models appear on Bedrock, SageMaker, Azure and Oracle OCI. Model Vault offers dedicated instances from $4 an hour. Cloudflare runs only on its own network, offers an OpenAI-compatible embeddings endpoint, and limits customization to bring-your-own LoRA on smaller models. Cohere says Command A+ reaches 375 tokens per second in 4-bit form. Regulated RAG inside your network points to Cohere. Public-cloud apps on Workers point to Cloudflare.
What Cloudflare Workers AI and Cohere do
Cloudflare Workers AI
Workers AI is the serverless GPU inference service of Cloudflare, which was founded in 2009. It launched in September 2023 and reached general availability in April 2024. Models run on GPUs inside Cloudflare's own network and are called from a Worker through an AI binding or over REST, including OpenAI-compatible Chat Completions and Embeddings endpoints plus a Responses endpoint for gpt-oss. The catalog lists 50+ open models. Since Kimi K2.5 arrived in March 2026 it has carried frontier-scale LLMs: Kimi K2.6 and K2.7 Code, GLM 5.2 and 5.3, DeepSeek V4 Pro and Flash, gpt-oss 120B and 20B, Qwen 3.8 27B and Llama 4 Scout. DeepSeek V4, added August 14, 2026, was the first to offer the full 1,048,576 token context.
Example models: DeepSeek V4 Pro, GLM 5.3
Full Cloudflare Workers AI profileCohere
Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.
Example models: Command A+, Command A, Embed 4, Rerank 4
Full Cohere profileShould you choose Cloudflare Workers AI or Cohere?
Cloudflare Workers AI
Choose Cloudflare Workers AI for
- Published per-token prices with no sales cycle
- Stronger open models for agentic coding
- Apps that already run on Cloudflare
Cohere
Choose Cohere for
- RAG and agents deployed inside a private VPC or on-prem
- Reranking and multimodal embeddings for search
- Multilingual assistants and translation
Cloudflare Workers AI vs Cohere at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Closed, plus open Command A+ |
| Flagship models | DeepSeek V4 Pro, GLM 5.3, Kimi K2.7 Code, gpt-oss 120B | Command A+, Command A, Embed 4, Rerank 4 |
| Speed | Unknown | 375 tok/s on Command A+ W4A4, per Cohere |
| Price | $0.011 per 1K Neurons; 10K free daily | $0.0375–$2.50 in, $0.15–$10 out per 1M |
| Customization | BYO LoRA on small models (beta) | Enterprise fine-tuning, incl. private |
| Deployment | Serverless on Cloudflare network | API, Bedrock, Azure, OCI, VPC, on-prem |
| Long context | 1M on DeepSeek V4; 262K on Kimi | 256K on Command A; 128K on A+ |
Frequently asked questions
What is the difference between Cloudflare Workers AI and Cohere?
Cohere builds for private enterprise RAG, with Embed, Rerank and on-prem deployment. Cloudflare Workers AI serves open models on shared edge infrastructure with public prices.
When should I choose Cloudflare Workers AI over Cohere?
Published per-token prices with no sales cycle; Stronger open models for agentic coding; Apps that already run on Cloudflare.
When should I choose Cohere over Cloudflare Workers AI?
RAG and agents deployed inside a private VPC or on-prem; Reranking and multimodal embeddings for search; Multilingual assistants and translation.
Is Cloudflare Workers AI or Cohere cheaper?
Cloudflare Workers AI: $0.011 per 1K Neurons; 10K free daily. Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. The cheaper choice depends on the model and workload.
Which has more context, Cloudflare Workers AI or Cohere?
Cloudflare Workers AI: 1M on DeepSeek V4; 262K on Kimi. Cohere: 256K on Command A; 128K on A+.
Related comparisons
Subconscious vs Cloudflare Workers AI
OpenAI vs Cloudflare Workers AI
Anthropic vs Cloudflare Workers AI
Google Vertex AI vs Cloudflare Workers AI
Amazon Bedrock vs Cloudflare Workers AI
Together AI vs Cloudflare Workers AI
Subconscious vs Cohere
OpenAI vs Cohere
Anthropic vs Cohere
Google Vertex AI vs Cohere
Amazon Bedrock vs Cohere
Together AI vs Cohere
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.