Long-running agents deserve better inference.
vs

Cerebras vs Luminal

Cerebras runs a few models on wafer-scale chips at record speed. Luminal compiles any model into faster code for GPUs and ASICs.

By The Subconscious Team · Updated

Cerebras vs Luminal: key differences

Cerebras is the fastest public host on what it serves, about 3,000 tokens per second on GPT-OSS 120B, but its self-serve catalog is two models and most others need a sales call. Luminal does not make chips. It compiles a model into native kernels for existing GPUs and ASICs, and reports 36K tokens per second aggregate on GPT-OSS 120B across 8 H100s.

Those numbers measure different things. Cerebras' figure is per-stream speed, which matters for voice and live UIs. Luminal's is total throughput across many requests, which matters for serving cost. Cerebras fits teams that need one stream as fast as possible; Luminal fits teams that want to run their own model efficiently on hardware they control.

What Cerebras and Luminal do

Cerebras

Cerebras builds a single chip the size of a silicon wafer, and its inference cloud is the fastest public host on the models it serves. Its developer table lists GPT-OSS 120B near 3,000 tokens per second, about six times Groq on the same weights, at $0.35 in and $0.75 out per million. The public shared catalog is thin, just GPT-OSS 120B and Gemma 4 31B as of August 2026. More model families live on dedicated endpoints and through partners like OpenRouter, Hugging Face, Vercel and AWS Marketplace.

Example models: GPT-OSS 120B, Gemma 4 31B

Full Cerebras profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Cerebras or Luminal?

Cerebras

Choose Cerebras for

  • Fastest per-stream output on supported models
  • Voice and live code completion
  • An ultrafast path to some OpenAI models

Luminal

Choose Luminal for

  • Serving custom or fine-tuned architectures off any catalog
  • Cutting serving cost per token on owned GPUs
  • An open-source engine teams can run on their own hardware

Cerebras vs Luminal at a glance

AttributeCerebrasLuminal
Model accessOpen weightsBring your own weights
Flagship modelsGPT-OSS 120B, Gemma 4 31BNo public catalog
Speed~3,000 tok/s on GPT-OSS 120B36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price$0.35 in, $0.75 out (GPT-OSS 120B)Pay per use; rates not published
CustomizationUnknownCompiles any PyTorch or HF model
DeploymentShared API, dedicated, partnersServerless (early access), on-prem license
Long contextUnknownUnknown

Frequently asked questions

What is the difference between Cerebras and Luminal?

Cerebras runs a few models on wafer-scale chips at record speed. Luminal compiles any model into faster code for GPUs and ASICs.

When should I choose Cerebras over Luminal?

Fastest per-stream output on supported models; Voice and live code completion; An ultrafast path to some OpenAI models.

When should I choose Luminal over Cerebras?

Serving custom or fine-tuned architectures off any catalog; Cutting serving cost per token on owned GPUs; An open-source engine teams can run on their own hardware.

Is Cerebras or Luminal cheaper?

Cerebras: $0.35 in, $0.75 out (GPT-OSS 120B). Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.