Modal vs Cohere
Modal rents per-second GPUs for code you write. Cohere sells finished models and retrieval tools, with private deployment for enterprises.
By The Subconscious Team · Updated
Modal vs Cohere: key differences
Modal sells compute, not models. A developer decorates a Python function with the GPU it needs, and Modal builds, scales and bills it per second, from $0.59 an hour for a T4 to $3.95 for an H100 at list. There is no catalog or per-token price, so teams bring their own weights and serving code. Cohere is the opposite: finished models behind an API. Command A lists at $2.50 in and $10 out with 256K context, Command R7B at $0.0375 in, and Command A+ is open under Apache 2.0, small enough to run on two H100s in 4-bit form, which means it could run on Modal too.
The choice is build versus buy. Modal suits spiky custom work like embeddings, transcription, fine-tunes and batch jobs, and its free Starter plan renews $30 of credits each month. Non-preemptible US production runs about 3.75x list, though, and keeping containers warm erodes the serverless savings. Cohere hands over Embed 4, Rerank 4, Aya and North ready to use, with fine-tuning and on-prem deployment for regulated buyers. Teams with ML engineers and custom models lean Modal. Teams that want a supported enterprise search stack without running GPUs lean Cohere.
What Modal and Cohere do
Modal
Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.
Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper
Full Modal profileCohere
Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.
Example models: Command A+, Command A, Embed 4, Rerank 4
Full Cohere profileShould you choose Modal or Cohere?
Modal vs Cohere at a glance
| Attribute | ||
|---|---|---|
| Model access | Bring your own weights | Closed, plus open Command A+ |
| Flagship models | None hosted | Command A+, Command A, Embed 4, Rerank 4 |
| Speed | ~1s container boot | 375 tok/s on Command A+ W4A4, per Cohere |
| Price | Per second; H100 $3.95/hr list | $0.0375–$2.50 in, $0.15–$10 out per 1M |
| Customization | Run any training code | Enterprise fine-tuning, incl. private |
| Deployment | Serverless GPU containers | API, Bedrock, Azure, OCI, VPC, on-prem |
| Long context | Depends on the model you deploy | 256K on Command A; 128K on A+ |
Frequently asked questions
What is the difference between Modal and Cohere?
Modal rents per-second GPUs for code you write. Cohere sells finished models and retrieval tools, with private deployment for enterprises.
When should I choose Modal over Cohere?
Custom or fine-tuned models on per-second GPUs; Bursty batch, embedding and transcription jobs; Python teams that want full control of serving.
When should I choose Cohere over Modal?
Ready-made retrieval without running GPUs; Supported on-prem enterprise deployments; Agent building on the North platform.
Is Modal or Cohere cheaper?
Modal: Per second; H100 $3.95/hr list. Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. The cheaper choice depends on the model and workload.
Which has more context, Modal or Cohere?
Modal: Depends on the model you deploy. Cohere: 256K on Command A; 128K on A+.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.