Cohere vs Thinking Machines
Thinking Machines gives researchers Tinker for custom SFT and RL on open weights. Cohere gives enterprises managed fine-tuning and private deployment.
By The Subconscious Team · Updated
Cohere vs Thinking Machines: key differences
Both offer customization, aimed at different teams. Thinking Machines' Tinker exposes four low-level calls so developers write their own supervised or reinforcement learning loops on Qwen3.5, Kimi K2.6, GLM-5.3, gpt-oss and its own Inkling models, while it handles distributed GPU work. Training is LoRA only and billed per million tokens across prefill, sample and train. Cohere's enterprise fine-tuning is a managed service on Command models, and it can run inside a customer VPC or on-prem. Researchers who want RL control lean Tinker; enterprises that want a supported tuning path on private data lean Cohere.
Serving is where Cohere is more complete. Thinking Machines' beta serverless API covers only Inkling and Inkling-Small, with Inkling at $1.00 in and $4.05 out and up to 1M tokens of context, and checkpoint sampling is scoped to testing and low internal traffic. Cohere serves Command A at $2.50 in and $10 out with 256K context, plus Embed 4 and Rerank 4, on its API, Bedrock, Azure and OCI. Inkling accepts text, image and audio input under Apache 2.0, and Command A+ covers reasoning, vision, translation and tool use in one Apache 2.0 model. Production RAG belongs on Cohere; custom post-training research belongs on Tinker.
What Cohere and Thinking Machines do
Cohere
Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.
Example models: Command A+, Command A, Embed 4, Rerank 4
Full Cohere profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Cohere or Thinking Machines?
Cohere
Choose Cohere for
- Production RAG with managed serving
- Supported fine-tuning inside a VPC
- Enterprise search on Embed and Rerank
Thinking Machines
Choose Thinking Machines for
- Custom RL loops on open MoE models
- Training task-specific LoRA adapters
- Testing Inkling with 1M-token context
Cohere vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Closed, plus open Command A+ | Open weights |
| Flagship models | Command A+, Command A, Embed 4, Rerank 4 | Inkling, Inkling-Small |
| Speed | 375 tok/s on Command A+ W4A4, per Cohere | Unknown |
| Price | $0.0375–$2.50 in, $0.15–$10 out per 1M | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Enterprise fine-tuning, incl. private | LoRA SFT and RL via Tinker |
| Deployment | API, Bedrock, Azure, OCI, VPC, on-prem | Training API, beta serverless (Inkling only) |
| Long context | 256K on Command A; 128K on A+ | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Cohere and Thinking Machines?
Thinking Machines gives researchers Tinker for custom SFT and RL on open weights. Cohere gives enterprises managed fine-tuning and private deployment.
When should I choose Cohere over Thinking Machines?
Production RAG with managed serving; Supported fine-tuning inside a VPC; Enterprise search on Embed and Rerank.
When should I choose Thinking Machines over Cohere?
Custom RL loops on open MoE models; Training task-specific LoRA adapters; Testing Inkling with 1M-token context.
Is Cohere or Thinking Machines cheaper?
Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Cohere or Thinking Machines?
Cohere: 256K on Command A; 128K on A+. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Cohere
OpenAI vs Cohere
Anthropic vs Cohere
Google Vertex AI vs Cohere
Amazon Bedrock vs Cohere
Together AI vs Cohere
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.