# Baseten vs Groq

> Baseten is fastest to the first token and hosts custom models. Groq streams output faster on its own LPU chip but runs a small catalog with no fine-tune hosting.

Canonical: https://www.subconscious.dev/compare/baseten-vs-groq · By The Subconscious Team · Updated September 30, 2026

## How they compare

These two are fast in different places. Groq is a speed chip company: its LPU keeps weights in on-chip SRAM, and Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B, with tight gaps between median and tail latency. Baseten runs GPUs and posted the lowest time to first token on the Artificial Analysis board in August 2026, 0.49 seconds. Groq wins when the output stream is the wait. Baseten wins when the first token is. Catalog breadth tilts toward Baseten, which serves 13 curated models including DeepSeek V4, GLM 5.2 and Kimi K3. Groq's list is small, shrinking and capped around 131K context.

Custom models decide many of these evaluations. A private fine-tune, a speech model or an embedding model can go onto Baseten through Truss at per-minute GPU rates with scale to zero. Groq offers no fine-tuned model hosting at all. Groq pushes back on price, with small-model tokens near the market floor and cache and Batch discounts that stack, while a dedicated Baseten H100 costs about $6.50 an hour. Baseten also carries HIPAA, data residency and a 99.99% SLA. Groq's future investment is less clear since NVIDIA hired most of its staff in late 2025.

## What each one does

### Baseten

Baseten runs two products. Model APIs serve a curated set of 13 open models, including DeepSeek V4, GLM 5.2, Kimi K3 and gpt-oss 120B, over endpoints that speak both the OpenAI Chat Completions shape and the Anthropic Messages shape. That dual compatibility means an existing OpenAI or Claude SDK, or a coding agent, points at Baseten with a base URL change. Dedicated deployments take any model you package with the open-source Truss CLI and bill per GPU minute, with an H100 at about $6.50 an hour.

### Groq

Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.

## Which is best, and when

### Choose Baseten for

- Hosting private fine-tunes, speech or embedding models
- Regulated teams that need HIPAA, data residency or self-hosting
- Agents built on DeepSeek V4, GLM 5.2 or Kimi K3

### Choose Groq for

- Voice agents on GPT-OSS or Qwen 3.6 where streaming speed matters
- Cheap small-model calls with stacked cache and Batch discounts
- Strict SLAs that depend on predictable tail latency

## At a glance

| Attribute | Baseten | Groq |
|---|---|---|
| Model access | Open weights, 13 curated | Open weights |
| Flagship models | GLM 5.2, DeepSeek V4, Kimi K3, gpt-oss 120B | GPT-OSS 120B, Qwen 3.6 27B |
| Speed | 0.49s TTFT, lowest measured | 500–1,000 tok/s |
| Price | H100 about $6.50/hr dedicated | Near the floor on small models |
| Customization | Deploy any model with Truss | No fine-tuned model hosting |
| Deployment | Model APIs, dedicated, self-host | GroqCloud API |
| Long context | Varies by model | Around 131K max |

## FAQ

### What is the difference between Baseten and Groq?

Baseten is fastest to the first token and hosts custom models. Groq streams output faster on its own LPU chip but runs a small catalog with no fine-tune hosting.

### When should I choose Baseten over Groq?

Hosting private fine-tunes, speech or embedding models; Regulated teams that need HIPAA, data residency or self-hosting; Agents built on DeepSeek V4, GLM 5.2 or Kimi K3.

### When should I choose Groq over Baseten?

Voice agents on GPT-OSS or Qwen 3.6 where streaming speed matters; Cheap small-model calls with stacked cache and Batch discounts; Strict SLAs that depend on predictable tail latency.

### Is Baseten or Groq cheaper?

Baseten: H100 about $6.50/hr dedicated. Groq: Near the floor on small models. The cheaper choice depends on the model and workload.

### Which has more context, Baseten or Groq?

Baseten: Varies by model. Groq: Around 131K max.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Baseten](https://www.subconscious.dev/compare/subconscious-vs-baseten.md), [Subconscious vs Groq](https://www.subconscious.dev/compare/subconscious-vs-groq.md).

Full profiles: [Baseten](https://www.subconscious.dev/providers/baseten.md), [Groq](https://www.subconscious.dev/providers/groq.md).
