Cohere vs Runware
Runware is a low-cost media generation API across image, video, audio and 3D. Cohere serves enterprise language and retrieval models, so overlap is thin.
By The Subconscious Team · Updated
Cohere vs Runware: key differences
Runware and Cohere serve different parts of a product. Runware's single request schema covers image, video, audio, 3D and text, with 300+ priced models, images from fractions of a cent and video by the second, such as Seedance 2.5 at about $0.10 a second at 480p. Its Sonic Inference Engine and Model Lake keep 400K+ models resident, and it rents H100s at $2.76 an hour. Cohere's catalog is language and search: Command A at $2.50 in and $10 out with 256K context, Command A+, Embed 4 for text, images and PDFs, Rerank 4, Aya and Transcribe.
Runware treats LLM hosting as a side line, so text workloads fit better with a dedicated language provider like Cohere. In the other direction, Cohere has no image or video generation. Deployment differs too. Runware is a hosted API plus raw GPUs, and output URLs expire after seven days by default, so apps need their own storage. Cohere offers private VPC and on-prem deployment with fine-tuning and sells through Bedrock, Azure and OCI. Runware supports fine-tuned diffusion checkpoints and community models at scale. A consumer app generating visuals picks Runware; an enterprise knowledge assistant picks Cohere.
What Cohere and Runware do
Cohere
Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.
Example models: Command A+, Command A, Embed 4, Rerank 4
Full Cohere profileRunware
Runware sells what it calls the lowest-cost API for media generation, and it claims more than 1M developers. One endpoint covers image, video, audio, 3D and text. Every request is a task with the same shape, so switching from a Kling video to a Seedream image mostly means changing the model ID. The published rate sheet lists 300+ priced models, with images from fractions of a cent to a few cents each and video billed per second, like Seedance 2.5 at about $0.10 a second at 480p.
Example models: Seedance 2.5, Qwen-Image-3.0
Full Runware profileShould you choose Cohere or Runware?
Cohere vs Runware at a glance
| Attribute | ||
|---|---|---|
| Model access | Closed, plus open Command A+ | Hosted media models |
| Flagship models | Command A+, Command A, Embed 4, Rerank 4 | Seedance 2.5, Qwen-Image-3.0 |
| Speed | 375 tok/s on Command A+ W4A4, per Cohere | Unknown |
| Price | $0.0375–$2.50 in, $0.15–$10 out per 1M | Images from fractions of a cent |
| Customization | Enterprise fine-tuning, incl. private | Fine-tuned diffusion checkpoints |
| Deployment | API, Bedrock, Azure, OCI, VPC, on-prem | Unified API, raw GPUs |
| Long context | 256K on Command A; 128K on A+ | Not applicable |
Frequently asked questions
What is the difference between Cohere and Runware?
Runware is a low-cost media generation API across image, video, audio and 3D. Cohere serves enterprise language and retrieval models, so overlap is thin.
When should I choose Cohere over Runware?
Enterprise knowledge assistants; Embedding PDFs and images for search; Private cloud language workloads.
When should I choose Runware over Cohere?
High-volume image generation at low cost; Short video clips billed per second; Serving community diffusion checkpoints.
Is Cohere or Runware cheaper?
Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. Runware: Images from fractions of a cent. The cheaper choice depends on the model and workload.
Which has more context, Cohere or Runware?
Cohere: 256K on Command A; 128K on A+. Runware: Not applicable.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.