# Cohere vs Thinking Machines

> Thinking Machines gives researchers Tinker for custom SFT and RL on open weights. Cohere gives enterprises managed fine-tuning and private deployment.

Canonical: https://www.subconscious.dev/compare/cohere-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Both offer customization, aimed at different teams. Thinking Machines' Tinker exposes four low-level calls so developers write their own supervised or reinforcement learning loops on Qwen3.5, Kimi K2.6, GLM-5.3, gpt-oss and its own Inkling models, while it handles distributed GPU work. Training is LoRA only and billed per million tokens across prefill, sample and train. Cohere's enterprise fine-tuning is a managed service on Command models, and it can run inside a customer VPC or on-prem. Researchers who want RL control lean Tinker; enterprises that want a supported tuning path on private data lean Cohere.

Serving is where Cohere is more complete. Thinking Machines' beta serverless API covers only Inkling and Inkling-Small, with Inkling at $1.00 in and $4.05 out and up to 1M tokens of context, and checkpoint sampling is scoped to testing and low internal traffic. Cohere serves Command A at $2.50 in and $10 out with 256K context, plus Embed 4 and Rerank 4, on its API, Bedrock, Azure and OCI. Inkling accepts text, image and audio input under Apache 2.0, and Command A+ covers reasoning, vision, translation and tool use in one Apache 2.0 model. Production RAG belongs on Cohere; custom post-training research belongs on Tinker.

## What each one does

### Cohere

Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Cohere for

- Production RAG with managed serving
- Supported fine-tuning inside a VPC
- Enterprise search on Embed and Rerank

### Choose Thinking Machines for

- Custom RL loops on open MoE models
- Training task-specific LoRA adapters
- Testing Inkling with 1M-token context

## At a glance

| Attribute | Cohere | Thinking Machines |
|---|---|---|
| Model access | Closed, plus open Command A+ | Open weights |
| Flagship models | Command A+, Command A, Embed 4, Rerank 4 | Inkling, Inkling-Small |
| Speed | 375 tok/s on Command A+ W4A4, per Cohere | - |
| Price | $0.0375–$2.50 in, $0.15–$10 out per 1M | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Enterprise fine-tuning, incl. private | LoRA SFT and RL via Tinker |
| Deployment | API, Bedrock, Azure, OCI, VPC, on-prem | Training API, beta serverless (Inkling only) |
| Long context | 256K on Command A; 128K on A+ | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Cohere and Thinking Machines?

Thinking Machines gives researchers Tinker for custom SFT and RL on open weights. Cohere gives enterprises managed fine-tuning and private deployment.

### When should I choose Cohere over Thinking Machines?

Production RAG with managed serving; Supported fine-tuning inside a VPC; Enterprise search on Embed and Rerank.

### When should I choose Thinking Machines over Cohere?

Custom RL loops on open MoE models; Training task-specific LoRA adapters; Testing Inkling with 1M-token context.

### Is Cohere or Thinking Machines cheaper?

Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Cohere or Thinking Machines?

Cohere: 256K on Command A; 128K on A+. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Cohere](https://www.subconscious.dev/compare/subconscious-vs-cohere.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Cohere](https://www.subconscious.dev/providers/cohere.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
