vs

Cerebras vs Modal

Cerebras sells the fastest tokens on its own chip. Modal rents general GPUs by the second for any model you bring. Different jobs, often in one stack.

By The Subconscious Team · Updated

Cerebras vs Modal: key differences

Cerebras is a speed product. It serves open models on a wafer-scale chip, lists GPT-OSS 120B near 3,000 tokens per second, and sells shared, dedicated and partner access. Modal is a compute product with no model catalog and no per-token price. Developers decorate Python functions with the GPU they need, from a T4 at $0.59 an hour to an H100 at $3.95 at list, and Modal builds, scales and bills the container per second. The overlap is small. Either can run an LLM, but only Cerebras offers a chip built for fast decode, and only Modal runs arbitrary code on GPUs.

Modal is the better fit for embeddings, reranking, transcription, OCR, fine-tuning and agent sandboxes, anything that needs custom code or custom weights. Cerebras is the better fit when output speed is the user's wait, such as voice or live code autocomplete. Watch the edge cases. Modal's non-preemptible US production runs about 3.75x list, and warm containers turn serverless into always-on. Cerebras' shared catalog is only GPT-OSS 120B and Gemma 4 31B, and its speed matters little when an agent mostly waits on tools.

What Cerebras and Modal do

Cerebras

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

Example models: GPT-OSS 120B, Gemma 4 31B

Full Cerebras profile

Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper

Full Modal profile

Should you choose Cerebras or Modal?

Cerebras

Choose Cerebras for

  • Voice agents that need tokens as fast as possible
  • Long generated outputs on GPT-OSS 120B
  • Teams that want speed without running infrastructure

Modal

Choose Modal for

  • Custom models, fine-tunes and batch jobs in Python
  • Spiky GPU work that scales to zero
  • Agent sandboxes and media jobs next to inference

Cerebras vs Modal at a glance

AttributeCerebrasModal
Model accessOpen weightsBring your own weights
Flagship modelsGPT-OSS 120B, Gemma 4 31BNone hosted
Speed~3,000 tok/s on GPT-OSS 120B~1s container boot
Price$0.35 in, $0.75 out (GPT-OSS 120B)Per second; H100 $3.95/hr list
CustomizationUnknownRun any training code
DeploymentShared API, dedicated, partnersServerless GPU containers
Long contextUnknownDepends on the model you deploy

Frequently asked questions

What is the difference between Cerebras and Modal?

Cerebras sells the fastest tokens on its own chip. Modal rents general GPUs by the second for any model you bring. Different jobs, often in one stack.

When should I choose Cerebras over Modal?

Voice agents that need tokens as fast as possible; Long generated outputs on GPT-OSS 120B; Teams that want speed without running infrastructure.

When should I choose Modal over Cerebras?

Custom models, fine-tunes and batch jobs in Python; Spiky GPU work that scales to zero; Agent sandboxes and media jobs next to inference.

Is Cerebras or Modal cheaper?

Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Modal: Per second; H100 $3.95/hr list. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.