We raised $5.1M for long-running agents.
vs

Groq vs Thinking Machines

Groq is about serving speed on a small catalog, with no fine-tuned model hosting. Thinking Machines is about building fine-tuned models, with little serving.

By The Subconscious Team · Updated

Groq vs Thinking Machines: key differences

Each covers what the other lacks. Groq runs open models on its LPU chip and publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B, with tail latency close to the median. It adds Whisper for speech to text and Groq Compound for built-in search and code execution. But its catalog is small, context tops out around 131K, and it hosts no fine-tuned models. Thinking Machines' Tinker is built around fine-tuning: four low-level calls let teams run SFT or RL with LoRA adapters on bases including gpt-oss, Qwen3.5, Kimi K2.6 and GLM-5.3. GPT-OSS-20B training bills $0.18 prefill, $0.45 sample and $0.40 train per million tokens.

Context and serving maturity separate them further. Thinking Machines' Inkling models reach up to 1M tokens with native image and audio input, well past Groq's cap, but the serverless API covering them is still in beta, and checkpoint sampling is not meant for user traffic. Groq is production-ready for latency-bound work today, though its long-term outlook is uncertain after NVIDIA licensed the LPU and hired most of its staff. A voice agent that needs every reply fast belongs on Groq. A team that needs a model trained on its own data, or context beyond 131K, has to look elsewhere, and Tinker covers the training half.

What Groq and Thinking Machines do

Groq

Groq serves open models on its own chip, the LPU, which keeps model weights in on-chip SRAM and runs a deterministic schedule instead of waiting on GPU memory. Groq publishes 1,000 tokens per second on GPT-OSS 20B and 500 on GPT-OSS 120B. Latency stays tight between median and tail, which matters for strict SLAs. The API is OpenAI-compatible and also hosts Whisper for speech to text plus an agentic system called Groq Compound with built-in search and code execution.

Example models: GPT-OSS 120B, Qwen 3.6 27B

Full Groq profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Groq or Thinking Machines?

Groq

Choose Groq for

  • Voice agents pairing Whisper with fast replies
  • Tight tail latency under strict SLAs
  • Fast multi-step loops on GPT-OSS models

Thinking Machines

Choose Thinking Machines for

  • Fine-tuning gpt-oss or Qwen3.5 on proprietary data
  • Open models with context past 131K
  • RL experiments with checkpoint sampling

Groq vs Thinking Machines at a glance

AttributeGroqThinking Machines
Model accessOpen weightsOpen weights
Flagship modelsGPT-OSS 120B, Qwen 3.6 27BInkling, Inkling-Small
Speed500–1,000 tok/sUnknown
PriceNear the floor on small modelsPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationNo fine-tuned model hostingLoRA SFT and RL via Tinker
DeploymentGroqCloud APITraining API, beta serverless (Inkling only)
Long contextAround 131K maxInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Groq and Thinking Machines?

Groq is about serving speed on a small catalog, with no fine-tuned model hosting. Thinking Machines is about building fine-tuned models, with little serving.

When should I choose Groq over Thinking Machines?

Voice agents pairing Whisper with fast replies; Tight tail latency under strict SLAs; Fast multi-step loops on GPT-OSS models.

When should I choose Thinking Machines over Groq?

Fine-tuning gpt-oss or Qwen3.5 on proprietary data; Open models with context past 131K; RL experiments with checkpoint sampling.

Is Groq or Thinking Machines cheaper?

Groq: Near the floor on small models. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Groq or Thinking Machines?

Groq: Around 131K max. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.