# Groq vs Cerebras

> Two custom-chip hosts chasing speed on open models. Cerebras is faster on GPT-OSS 120B; Groq is cheaper per token, steadier at the tail and carries a few more models.

Canonical: https://www.subconscious.dev/compare/groq-vs-cerebras · By The Subconscious Team · Updated September 30, 2026

## How they compare

Groq and Cerebras are the two chip companies most developers compare for raw speed. Cerebras's wafer-scale chip lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq's published 500 on the same weights. Groq answers with price and predictability. Its per-token rates on small models sit near the market floor, cache and Batch discounts stack, and its deterministic LPU schedule keeps median and tail latency close. Cerebras charges $0.35 in and $0.75 out on GPT-OSS 120B and, by its own account, costs more per token than Groq on shared models. Both catalogs are thin. Cerebras's shared list is GPT-OSS 120B and Gemma 4 31B; Groq's centers on GPT-OSS and Qwen 3.6.

Extras and outlook break the tie. Groq hosts Whisper for speech to text and Groq Compound, an agentic system with built-in search and code execution, which suits voice agents. Cerebras reaches more models through dedicated endpoints and partners like OpenRouter and AWS Marketplace, and it runs OpenAI's Ultrafast GPT-5.6 Sol preview. Groq's future is murkier since NVIDIA licensed the LPU and hired most of its engineers. Cerebras now trades publicly on Nasdaq.

## What each one does

### Groq

Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.

### Cerebras

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

## Which is best, and when

### Choose Groq for

- Voice agents pairing Whisper with fast LLM replies
- Cost-sensitive small-model calls with stacked discounts
- SLAs judged on tail latency rather than peak speed

### Choose Cerebras for

- Maximum tokens per second on GPT-OSS 120B
- Long streamed outputs in live code tools
- Access to GPT-5.6 Sol Ultrafast on wafer-scale hardware

## At a glance

| Attribute | Groq | Cerebras |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GPT-OSS 120B, Qwen 3.6 27B | GPT-OSS 120B, Gemma 4 31B |
| Speed | 500–1,000 tok/s | ~3,000 tok/s on GPT-OSS 120B |
| Price | Near the floor on small models | $0.35 in, $0.75 out (GPT-OSS 120B) |
| Customization | No fine-tuned model hosting | - |
| Deployment | GroqCloud API | Shared API, dedicated, partners |
| Long context | Around 131K max | - |

## FAQ

### What is the difference between Groq and Cerebras?

Two custom-chip hosts chasing speed on open models. Cerebras is faster on GPT-OSS 120B; Groq is cheaper per token, steadier at the tail and carries a few more models.

### When should I choose Groq over Cerebras?

Voice agents pairing Whisper with fast LLM replies; Cost-sensitive small-model calls with stacked discounts; SLAs judged on tail latency rather than peak speed.

### When should I choose Cerebras over Groq?

Maximum tokens per second on GPT-OSS 120B; Long streamed outputs in live code tools; Access to GPT-5.6 Sol Ultrafast on wafer-scale hardware.

### Is Groq or Cerebras cheaper?

Groq: Near the floor on small models. Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Groq](https://www.subconscious.dev/compare/subconscious-vs-groq.md), [Subconscious vs Cerebras](https://www.subconscious.dev/compare/subconscious-vs-cerebras.md).

Full profiles: [Groq](https://www.subconscious.dev/providers/groq.md), [Cerebras](https://www.subconscious.dev/providers/cerebras.md).
