# Groq

> Extreme speed on its own LPU chip, with tight tail latency on a small catalog.

Canonical: https://www.subconscious.dev/providers/groq · By The Subconscious Team · Updated September 30, 2026

- Founded: 2016
- Example models: GPT-OSS 120B, Qwen 3.6 27B
- Website: https://groq.com

## Overview

Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.

The company changed shape in December 2025, when NVIDIA licensed the LPU design and hired roughly 90% of staff, founder Jonathan Ross included, in a deal reported near $20B. GroqCloud keeps operating independently under CEO Simon Edwards, and NVIDIA showed a Groq 3 LPU inside Vera Rubin at GTC 2026. A Senate antitrust inquiry into the deal structure opened in March 2026. The catalog has also narrowed: Llama 3.3 70B and Llama 3.1 8B shut down on August 16, 2026, with Groq steering users to GPT-OSS and Qwen 3.6.

## Upsides

- Several times the tokens per second of GPU hosts on supported models.
- Predictable tail latency.
- Per-token prices on its small models sit near the market floor, with cache and Batch discounts that stack.

## Core use cases

- Voice agents where any pause reads as awkward.
- Multi-call agent loops where each step has to return fast.

## Downsides

- Small, shrinking catalog capped around 131K context, with no fine-tuned model hosting.
- Long-term investment in GroqCloud is an open question now that its core engineers work at NVIDIA.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open weights |
| Flagship models | GPT-OSS 120B, Qwen 3.6 27B |
| Speed | 500–1,000 tok/s |
| Price | Near the floor on small models |
| Customization | No fine-tuned model hosting |
| Deployment | GroqCloud API |
| Long context | Around 131K max |

## FAQ

### What is Groq?

Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.

### What is Groq best for?

Voice agents where any pause reads as awkward; Multi-call agent loops where each step has to return fast.

### How much does Groq cost?

Groq pricing at a glance: Near the floor on small models. Rates change often, so check Groq's pricing page before committing.

### How much context does Groq support?

Groq's long-context support: Around 131K max.

### What are the downsides of Groq?

Small, shrinking catalog capped around 131K context, with no fine-tuned model hosting; Long-term investment in GroqCloud is an open question now that its core engineers work at NVIDIA.

### What are the best alternatives to Groq?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Groq on this site.

## Comparisons

- [Subconscious vs Groq](https://www.subconscious.dev/compare/subconscious-vs-groq.md)
- [OpenAI vs Groq](https://www.subconscious.dev/compare/openai-vs-groq.md)
- [Anthropic vs Groq](https://www.subconscious.dev/compare/anthropic-vs-groq.md)
- [Google Vertex AI vs Groq](https://www.subconscious.dev/compare/google-vertex-vs-groq.md)
- [Amazon Bedrock vs Groq](https://www.subconscious.dev/compare/aws-bedrock-vs-groq.md)
- [Together AI vs Groq](https://www.subconscious.dev/compare/together-ai-vs-groq.md)
- [Fireworks AI vs Groq](https://www.subconscious.dev/compare/fireworks-vs-groq.md)
- [Baseten vs Groq](https://www.subconscious.dev/compare/baseten-vs-groq.md)
- [Groq vs Cerebras](https://www.subconscious.dev/compare/groq-vs-cerebras.md)
- [Groq vs DeepInfra](https://www.subconscious.dev/compare/groq-vs-deepinfra.md)
- [Groq vs Modal](https://www.subconscious.dev/compare/groq-vs-modal.md)
- [Groq vs xAI](https://www.subconscious.dev/compare/groq-vs-xai.md)
- [Groq vs DeepSeek](https://www.subconscious.dev/compare/groq-vs-deepseek.md)
- [Groq vs Moonshot AI](https://www.subconscious.dev/compare/groq-vs-moonshot-ai.md)
- [Groq vs Z.ai](https://www.subconscious.dev/compare/groq-vs-z-ai.md)
- [Groq vs Alibaba Cloud](https://www.subconscious.dev/compare/groq-vs-alibaba-cloud.md)
- [Groq vs Meta](https://www.subconscious.dev/compare/groq-vs-meta.md)
- [Groq vs SambaNova](https://www.subconscious.dev/compare/groq-vs-sambanova.md)
- [Groq vs Nebius](https://www.subconscious.dev/compare/groq-vs-nebius.md)
- [Groq vs fal](https://www.subconscious.dev/compare/groq-vs-fal.md)
- [Groq vs Novita AI](https://www.subconscious.dev/compare/groq-vs-novita-ai.md)
- [Groq vs Parasail](https://www.subconscious.dev/compare/groq-vs-parasail.md)
- [Groq vs Inference.net](https://www.subconscious.dev/compare/groq-vs-inference-net.md)
- [Groq vs GMI Cloud](https://www.subconscious.dev/compare/groq-vs-gmi-cloud.md)
- [Groq vs Sail Research](https://www.subconscious.dev/compare/groq-vs-sail-research.md)
- [Groq vs Morph](https://www.subconscious.dev/compare/groq-vs-morph.md)
- [Groq vs Relace](https://www.subconscious.dev/compare/groq-vs-relace.md)
- [Groq vs TypeSafe AI](https://www.subconscious.dev/compare/groq-vs-typesafe-ai.md)
- [Groq vs StepFun](https://www.subconscious.dev/compare/groq-vs-stepfun.md)
- [Groq vs Runware](https://www.subconscious.dev/compare/groq-vs-runware.md)
- [Groq vs StreamLake](https://www.subconscious.dev/compare/groq-vs-streamlake.md)
- [Groq vs Wafer](https://www.subconscious.dev/compare/groq-vs-wafer.md)
- [Groq vs RunInfra](https://www.subconscious.dev/compare/groq-vs-runinfra.md)
- [Groq vs Particle.AI](https://www.subconscious.dev/compare/groq-vs-particle-ai.md)

## Sources

- [Groq benchmark 2026, Markaicode](https://markaicode.com/benchmarks/groq-production-benchmark-latency/)
- [Groq explained, Layer3](https://www.layer3labs.io/guides/groq-explained)

Pricing and model lineups change often; figures are a snapshot.
