# Cerebras vs Modal

> Cerebras sells the fastest tokens on its own chip. Modal rents general GPUs by the second for any model you bring. Different jobs, often in one stack.

Canonical: https://www.subconscious.dev/compare/cerebras-vs-modal · By The Subconscious Team · Updated September 30, 2026

## How they compare

Cerebras is a speed product. It serves open models on a wafer-scale chip, lists GPT-OSS 120B near 3,000 tokens per second, and sells shared, dedicated and partner access. Modal is a compute product with no model catalog and no per-token price. Developers decorate Python functions with the GPU they need, from a T4 at $0.59 an hour to an H100 at $3.95 at list, and Modal builds, scales and bills the container per second. The overlap is small. Either can run an LLM, but only Cerebras offers a chip built for fast decode, and only Modal runs arbitrary code on GPUs.

Modal is the better fit for embeddings, reranking, transcription, OCR, fine-tuning and agent sandboxes, anything that needs custom code or custom weights. Cerebras is the better fit when output speed is the user's wait, such as voice or live code autocomplete. Watch the edge cases. Modal's non-preemptible US production runs about 3.75x list, and warm containers turn serverless into always-on. Cerebras' shared catalog is only GPT-OSS 120B and Gemma 4 31B, and its speed matters little when an agent mostly waits on tools.

## What each one does

### Cerebras

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

## Which is best, and when

### Choose Cerebras for

- Voice agents that need tokens as fast as possible
- Long generated outputs on GPT-OSS 120B
- Teams that want speed without running infrastructure

### Choose Modal for

- Custom models, fine-tunes and batch jobs in Python
- Spiky GPU work that scales to zero
- Agent sandboxes and media jobs next to inference

## At a glance

| Attribute | Cerebras | Modal |
|---|---|---|
| Model access | Open weights | Bring your own weights |
| Flagship models | GPT-OSS 120B, Gemma 4 31B | None hosted |
| Speed | ~3,000 tok/s on GPT-OSS 120B | ~1s container boot |
| Price | $0.35 in, $0.75 out (GPT-OSS 120B) | Per second; H100 $3.95/hr list |
| Customization | - | Run any training code |
| Deployment | Shared API, dedicated, partners | Serverless GPU containers |
| Long context | - | Depends on the model you deploy |

## FAQ

### What is the difference between Cerebras and Modal?

Cerebras sells the fastest tokens on its own chip. Modal rents general GPUs by the second for any model you bring. Different jobs, often in one stack.

### When should I choose Cerebras over Modal?

Voice agents that need tokens as fast as possible; Long generated outputs on GPT-OSS 120B; Teams that want speed without running infrastructure.

### When should I choose Modal over Cerebras?

Custom models, fine-tunes and batch jobs in Python; Spiky GPU work that scales to zero; Agent sandboxes and media jobs next to inference.

### Is Cerebras or Modal cheaper?

Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Modal: Per second; H100 $3.95/hr list. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Cerebras](https://www.subconscious.dev/compare/subconscious-vs-cerebras.md), [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md).

Full profiles: [Cerebras](https://www.subconscious.dev/providers/cerebras.md), [Modal](https://www.subconscious.dev/providers/modal.md).
