# Groq vs Thinking Machines

> Groq is about serving speed on a small catalog, with no fine-tuned model hosting. Thinking Machines is about building fine-tuned models, with little serving.

Canonical: https://www.subconscious.dev/compare/groq-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Each covers what the other lacks. Groq runs open models on its LPU chip and publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B, with tail latency close to the median. It adds Whisper for speech to text and Groq Compound for built-in search and code execution. But its catalog is small, context tops out around 131K, and it hosts no fine-tuned models. Thinking Machines' Tinker is built around fine-tuning: four low-level calls let teams run SFT or RL with LoRA adapters on bases including gpt-oss, Qwen3.5, Kimi K2.6 and GLM-5.3. GPT-OSS-20B training bills $0.18 prefill, $0.45 sample and $0.40 train per million tokens.

Context and serving maturity separate them further. Thinking Machines' Inkling models reach up to 1M tokens with native image and audio input, well past Groq's cap, but the serverless API covering them is still in beta, and checkpoint sampling is not meant for user traffic. Groq is production-ready for latency-bound work today, though its long-term outlook is uncertain after NVIDIA licensed the LPU and hired most of its staff. A voice agent that needs every reply fast belongs on Groq. A team that needs a model trained on its own data, or context beyond 131K, has to look elsewhere, and Tinker covers the training half.

## What each one does

### Groq

Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Groq for

- Voice agents pairing Whisper with fast replies
- Tight tail latency under strict SLAs
- Fast multi-step loops on GPT-OSS models

### Choose Thinking Machines for

- Fine-tuning gpt-oss or Qwen3.5 on proprietary data
- Open models with context past 131K
- RL experiments with checkpoint sampling

## At a glance

| Attribute | Groq | Thinking Machines |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GPT-OSS 120B, Qwen 3.6 27B | Inkling, Inkling-Small |
| Speed | 500–1,000 tok/s | - |
| Price | Near the floor on small models | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | No fine-tuned model hosting | LoRA SFT and RL via Tinker |
| Deployment | GroqCloud API | Training API, beta serverless (Inkling only) |
| Long context | Around 131K max | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Groq and Thinking Machines?

Groq is about serving speed on a small catalog, with no fine-tuned model hosting. Thinking Machines is about building fine-tuned models, with little serving.

### When should I choose Groq over Thinking Machines?

Voice agents pairing Whisper with fast replies; Tight tail latency under strict SLAs; Fast multi-step loops on GPT-OSS models.

### When should I choose Thinking Machines over Groq?

Fine-tuning gpt-oss or Qwen3.5 on proprietary data; Open models with context past 131K; RL experiments with checkpoint sampling.

### Is Groq or Thinking Machines cheaper?

Groq: Near the floor on small models. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Groq or Thinking Machines?

Groq: Around 131K max. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Groq](https://www.subconscious.dev/compare/subconscious-vs-groq.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Groq](https://www.subconscious.dev/providers/groq.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
