# Cerebras vs Cohere

> Cerebras sells record decode speed on a thin open-model catalog. Cohere sells its own enterprise models with retrieval tools and private deployment.

Canonical: https://www.subconscious.dev/compare/cerebras-vs-cohere · By The Subconscious Team · Updated September 30, 2026

## How they compare

Cerebras is the fastest public host on the models it runs, listing GPT-OSS 120B near 3,000 tokens per second at $0.35 in and $0.75 out. Its shared catalog is just GPT-OSS 120B and Gemma 4 31B, with more on dedicated endpoints and partners like OpenRouter and AWS Marketplace, and it powers OpenAI's Ultrafast GPT-5.6 Sol preview. Cohere claims 375 tokens per second on Command A+ in 4-bit form. Cohere's Command A lists at $2.50 in and $10 out with 256K context, so on raw speed and price per token Cerebras's shared models are hard to beat.

Cohere wins nearly everywhere else an enterprise buyer looks. Cerebras lists no fine-tuning, while Cohere fine-tunes inside private environments and installs in any VPC or on-prem. Embed 4, Rerank 4, Aya and Transcribe cover retrieval, multilingual and speech. Command A+ is open under Apache 2.0 and runs on two H100s or one B200 in 4-bit form, so it does not need exotic hardware. Choose Cerebras for streaming UIs and live code tools where generation speed is the wait. Choose Cohere when the job is grounded answers over private documents.

## What each one does

### Cerebras

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

### Cohere

Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.

## Which is best, and when

### Choose Cerebras for

- Maximum tokens per second on GPT-OSS 120B
- Streaming UIs and live code autocomplete
- Ultrafast GPT-5.6 Sol access

### Choose Cohere for

- Grounded answers over private documents
- Self-hosting on two H100s under Apache 2.0
- Multilingual and speech models in one vendor

## At a glance

| Attribute | Cerebras | Cohere |
|---|---|---|
| Model access | Open weights | Closed, plus open Command A+ |
| Flagship models | GPT-OSS 120B, Gemma 4 31B | Command A+, Command A, Embed 4, Rerank 4 |
| Speed | ~3,000 tok/s on GPT-OSS 120B | 375 tok/s on Command A+ W4A4, per Cohere |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) | $0.0375–$2.50 in, $0.15–$10 out per 1M |
| Customization | - | Enterprise fine-tuning, incl. private |
| Deployment | Shared API, dedicated, partners | API, Bedrock, Azure, OCI, VPC, on-prem |
| Long context | - | 256K on Command A; 128K on A+ |

## FAQ

### What is the difference between Cerebras and Cohere?

Cerebras sells record decode speed on a thin open-model catalog. Cohere sells its own enterprise models with retrieval tools and private deployment.

### When should I choose Cerebras over Cohere?

Maximum tokens per second on GPT-OSS 120B; Streaming UIs and live code autocomplete; Ultrafast GPT-5.6 Sol access.

### When should I choose Cohere over Cerebras?

Grounded answers over private documents; Self-hosting on two H100s under Apache 2.0; Multilingual and speech models in one vendor.

### Is Cerebras or Cohere cheaper?

Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Cerebras](https://www.subconscious.dev/compare/subconscious-vs-cerebras.md), [Subconscious vs Cohere](https://www.subconscious.dev/compare/subconscious-vs-cohere.md).

Full profiles: [Cerebras](https://www.subconscious.dev/providers/cerebras.md), [Cohere](https://www.subconscious.dev/providers/cohere.md).
