Groq vs Cohere
Groq serves a few open models very fast on its LPU. Cohere serves its own enterprise models for RAG, with private and on-prem installs Groq does not offer.
By The Subconscious Team · Updated
Groq vs Cohere: key differences
Groq optimizes for speed. Its LPU publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B, with tight tail latency, and its small-model prices sit near the market floor. The catalog is small and capped around 131K context, it hosts no fine-tuned models, and it runs only on GroqCloud. Cohere's Command A has a 256K window at $2.50 in and $10 out, Command R7B costs $0.0375 in, and Command A+ runs 128K. Cohere claims 375 tokens per second on Command A+ in 4-bit form, which is fast for a 218B model but below Groq's published figures on smaller models.
Cohere wins on deployment and customization. It fine-tunes inside customer environments, installs in any VPC or on-prem, and sells through Bedrock, Azure AI Foundry and OCI. Embed 4 and Rerank 4 give it retrieval models Groq lacks, while Groq offers Whisper and Groq Compound for voice and agentic search. Groq's long-term outlook is uncertain after NVIDIA licensed the LPU and hired most of its staff. Command A+ weights are also open under Apache 2.0, so teams can self-host rather than rely on one cloud. For voice agents and tight latency budgets, Groq. For private enterprise search and grounded answers, Cohere.
What Groq and Cohere do
Groq
Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.
Example models: GPT-OSS 120B, Qwen 3.6 27B
Full Groq profileCohere
Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.
Example models: Command A+, Command A, Embed 4, Rerank 4
Full Cohere profileShould you choose Groq or Cohere?
Groq vs Cohere at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Closed, plus open Command A+ |
| Flagship models | GPT-OSS 120B, Qwen 3.6 27B | Command A+, Command A, Embed 4, Rerank 4 |
| Speed | 500–1,000 tok/s | 375 tok/s on Command A+ W4A4, per Cohere |
| Price | Near the floor on small models | $0.0375–$2.50 in, $0.15–$10 out per 1M |
| Customization | No fine-tuned model hosting | Enterprise fine-tuning, incl. private |
| Deployment | GroqCloud API | API, Bedrock, Azure, OCI, VPC, on-prem |
| Long context | Around 131K max | 256K on Command A; 128K on A+ |
Frequently asked questions
What is the difference between Groq and Cohere?
Groq serves a few open models very fast on its LPU. Cohere serves its own enterprise models for RAG, with private and on-prem installs Groq does not offer.
When should I choose Groq over Cohere?
Voice agents that need instant replies; Strict tail-latency SLAs; Cheap small-model calls at high volume.
When should I choose Cohere over Groq?
Private RAG inside a customer network; Fine-tuned enterprise models; Reranking and embeddings for search.
Is Groq or Cohere cheaper?
Groq: Near the floor on small models. Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. The cheaper choice depends on the model and workload.
Which has more context, Groq or Cohere?
Groq: Around 131K max. Cohere: 256K on Command A; 128K on A+.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.