# Cerebras

> The fastest public inference host, running models on a wafer-scale chip.

Canonical: https://www.subconscious.dev/providers/cerebras · By The Subconscious Team · Updated September 30, 2026

- Founded: 2015
- Example models: GPT-OSS 120B, Gemma 4 31B
- Website: https://www.cerebras.ai

## Overview

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

OpenAI is the anchor customer. A January 2026 deal worth over $10B has OpenAI renting roughly 750 MW of Cerebras capacity through 2028, and on August 13, 2026 OpenAI previewed an Ultrafast tier of GPT-5.6 Sol running on Cerebras at up to 750 output tokens per second. Cerebras now trades on Nasdaq as CBRS. In late September it announced CS-4 hardware and a Gimlet Labs partnership that pairs its wafers with GPUs in a disaggregated cloud targeting 3,000 tokens per second.

## Upsides

- Fastest published tokens per second of any public inference host.
- The only wafer-scale path to a closed frontier model, via OpenAI's Ultrafast preview.

## Core use cases

- Voice, live code autocomplete and streaming UIs where generation is the wait.
- Agent steps that emit long outputs on open weights.

## Downsides

- Tiny self-serve catalog, so most models mean a sales conversation.
- Costs more per token than Groq on shared models, and the speed does little when an agent mostly waits on tools or hidden reasoning.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open weights |
| Flagship models | GPT-OSS 120B, Gemma 4 31B |
| Speed | ~3,000 tok/s on GPT-OSS 120B |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) |
| Customization | - |
| Deployment | Shared API, dedicated, partners |
| Long context | - |

## FAQ

### What is Cerebras?

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

### What is Cerebras best for?

Voice, live code autocomplete and streaming UIs where generation is the wait; Agent steps that emit long outputs on open weights.

### How much does Cerebras cost?

Cerebras pricing at a glance: $0.35 in, $0.75 out (GPT-OSS 120B). Rates change often, so check Cerebras's pricing page before committing.

### What are the downsides of Cerebras?

Tiny self-serve catalog, so most models mean a sales conversation; Costs more per token than Groq on shared models, and the speed does little when an agent mostly waits on tools or hidden reasoning.

### What are the best alternatives to Cerebras?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Cerebras on this site.

## Comparisons

- [Subconscious vs Cerebras](https://www.subconscious.dev/compare/subconscious-vs-cerebras.md)
- [OpenAI vs Cerebras](https://www.subconscious.dev/compare/openai-vs-cerebras.md)
- [Anthropic vs Cerebras](https://www.subconscious.dev/compare/anthropic-vs-cerebras.md)
- [Google Vertex AI vs Cerebras](https://www.subconscious.dev/compare/google-vertex-vs-cerebras.md)
- [Amazon Bedrock vs Cerebras](https://www.subconscious.dev/compare/aws-bedrock-vs-cerebras.md)
- [Together AI vs Cerebras](https://www.subconscious.dev/compare/together-ai-vs-cerebras.md)
- [Fireworks AI vs Cerebras](https://www.subconscious.dev/compare/fireworks-vs-cerebras.md)
- [Baseten vs Cerebras](https://www.subconscious.dev/compare/baseten-vs-cerebras.md)
- [Groq vs Cerebras](https://www.subconscious.dev/compare/groq-vs-cerebras.md)
- [Cerebras vs DeepInfra](https://www.subconscious.dev/compare/cerebras-vs-deepinfra.md)
- [Cerebras vs Modal](https://www.subconscious.dev/compare/cerebras-vs-modal.md)
- [Cerebras vs xAI](https://www.subconscious.dev/compare/cerebras-vs-xai.md)
- [Cerebras vs DeepSeek](https://www.subconscious.dev/compare/cerebras-vs-deepseek.md)
- [Cerebras vs Moonshot AI](https://www.subconscious.dev/compare/cerebras-vs-moonshot-ai.md)
- [Cerebras vs Z.ai](https://www.subconscious.dev/compare/cerebras-vs-z-ai.md)
- [Cerebras vs Alibaba Cloud](https://www.subconscious.dev/compare/cerebras-vs-alibaba-cloud.md)
- [Cerebras vs Meta](https://www.subconscious.dev/compare/cerebras-vs-meta.md)
- [Cerebras vs SambaNova](https://www.subconscious.dev/compare/cerebras-vs-sambanova.md)
- [Cerebras vs Nebius](https://www.subconscious.dev/compare/cerebras-vs-nebius.md)
- [Cerebras vs fal](https://www.subconscious.dev/compare/cerebras-vs-fal.md)
- [Cerebras vs Novita AI](https://www.subconscious.dev/compare/cerebras-vs-novita-ai.md)
- [Cerebras vs Parasail](https://www.subconscious.dev/compare/cerebras-vs-parasail.md)
- [Cerebras vs Inference.net](https://www.subconscious.dev/compare/cerebras-vs-inference-net.md)
- [Cerebras vs GMI Cloud](https://www.subconscious.dev/compare/cerebras-vs-gmi-cloud.md)
- [Cerebras vs Sail Research](https://www.subconscious.dev/compare/cerebras-vs-sail-research.md)
- [Cerebras vs Morph](https://www.subconscious.dev/compare/cerebras-vs-morph.md)
- [Cerebras vs Relace](https://www.subconscious.dev/compare/cerebras-vs-relace.md)
- [Cerebras vs TypeSafe AI](https://www.subconscious.dev/compare/cerebras-vs-typesafe-ai.md)
- [Cerebras vs StepFun](https://www.subconscious.dev/compare/cerebras-vs-stepfun.md)
- [Cerebras vs Runware](https://www.subconscious.dev/compare/cerebras-vs-runware.md)
- [Cerebras vs StreamLake](https://www.subconscious.dev/compare/cerebras-vs-streamlake.md)
- [Cerebras vs Wafer](https://www.subconscious.dev/compare/cerebras-vs-wafer.md)
- [Cerebras vs RunInfra](https://www.subconscious.dev/compare/cerebras-vs-runinfra.md)
- [Cerebras vs Particle.AI](https://www.subconscious.dev/compare/cerebras-vs-particle-ai.md)

## Sources

- [Cerebras vs Groq, BenchLM](https://benchlm.ai/blog/posts/cerebras-won-speed-not-the-shortlist)
- [Cerebras OpenAI deal, The Next Platform](https://www.nextplatform.com/ai/2026/01/15/cerebras-inks-transformative-10-billion-inference-deal-with-openai/4092155)
- [Cerebras and Gimlet Labs](https://www.cerebras.ai/press-release/gimlet-labs-adds-cerebras-to-deliver-ultrafast-ai-inference-through-gimlet-cloud-deployment)

Pricing and model lineups change often; figures are a snapshot.
