Together AI vs Cohere
Together AI hosts 30+ open models with training and GPU clusters on one bill. Cohere sells its own Command, Embed and Rerank models to enterprises that want them in private networks.
By The Subconscious Team · Updated
Together AI vs Cohere: key differences
Together AI is a broad open-model host. Its text catalog runs past thirty models, including DeepSeek V4, Kimi K3, GLM 5.2 and Qwen 3.8, and new releases land within days. One bill covers serverless, batch at up to 50% off, dedicated endpoints and raw GPU clusters with H100s from $3.19 an hour reserved. Cohere sells its own lineup: Command A at $2.50 in and $10 out with 256K context, Command A+ under Apache 2.0 with 128K, and Command R7B at $0.0375 in. On model choice and agentic coding strength, Together's catalog has the edge, since Cohere says Command A+ trails the latest DeepSeek and GLM models there.
Both offer fine-tuning, but in different settings. Together runs LoRA and full SFT from $0.48 per million training tokens, with RL in closed beta, all on its own cloud. Cohere fine-tunes inside a customer's VPC or on-prem, which matters to banks and governments. Cohere also brings Embed 4 and Rerank 4, a retrieval stack Together does not center on. Together has no free tier, while Cohere's Command A+ prices are unpublished, so both involve cost discovery. Pick Together for open-model breadth and training, Cohere for private enterprise search.
What Together AI and Cohere do
Together AI
Together AI is the broadest open-model platform in the category. One bill covers per-token serverless inference, batch at up to 50% off, provisioned throughput with a 99% SLA, dedicated deployments, raw GPU clusters, managed fine-tuning and code sandboxes for agents. The text catalog runs past thirty open models, including DeepSeek V4, Kimi K3, GLM 5.2, Qwen 3.8 and MiniMax M3, plus image, video, speech and embedding models. Token prices sit at parity with Fireworks and Baseten.
Example models: Kimi K3, DeepSeek V4 Pro
Full Together AI profileCohere
Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.
Example models: Command A+, Command A, Embed 4, Rerank 4
Full Cohere profileShould you choose Together AI or Cohere?
Together AI
Choose Together AI for
- Choosing among 30+ open models on one API
- Fine-tuning then serving on the same platform
- Reserved GPU clusters for large experiments
Cohere
Choose Cohere for
- Fine-tuning inside a private network
- Rerank and multimodal embeddings for search
- Regulated buyers needing on-prem generation
Together AI vs Cohere at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Closed, plus open Command A+ |
| Flagship models | Kimi K3, DeepSeek V4, GLM 5.2, Qwen 3.8 | Command A+, Command A, Embed 4, Rerank 4 |
| Speed | 0.99s TTFT on DeepSeek V4 Pro | 375 tok/s on Command A+ W4A4, per Cohere |
| Price | Parity with Fireworks and Baseten | $0.0375–$2.50 in, $0.15–$10 out per 1M |
| Customization | LoRA and full SFT; RL in beta | Enterprise fine-tuning, incl. private |
| Deployment | Serverless, dedicated, GPU clusters | API, Bedrock, Azure, OCI, VPC, on-prem |
| Long context | 512K on DeepSeek V4 Pro | 256K on Command A; 128K on A+ |
Frequently asked questions
What is the difference between Together AI and Cohere?
Together AI hosts 30+ open models with training and GPU clusters on one bill. Cohere sells its own Command, Embed and Rerank models to enterprises that want them in private networks.
When should I choose Together AI over Cohere?
Choosing among 30+ open models on one API; Fine-tuning then serving on the same platform; Reserved GPU clusters for large experiments.
When should I choose Cohere over Together AI?
Fine-tuning inside a private network; Rerank and multimodal embeddings for search; Regulated buyers needing on-prem generation.
Is Together AI or Cohere cheaper?
Together AI: Parity with Fireworks and Baseten. Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. The cheaper choice depends on the model and workload.
Which has more context, Together AI or Cohere?
Together AI: 512K on DeepSeek V4 Pro. Cohere: 256K on Command A; 128K on A+.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.