Cohere vs Luminal
Cohere sells enterprise models for RAG and search that run in a VPC or on-prem. Luminal sells a compiler that speeds up models you choose.
By The Subconscious Team · Updated
Cohere vs Luminal: key differences
Cohere offers Command A+, Embed 4 and Rerank 4 for retrieval and agents, deployable through its API, the major clouds, a VPC or on-prem, with private fine-tuning. Luminal brings no models. Its compiler lowers a model you supply into fused GPU kernels ahead of time, served on early-access endpoints or under an on-prem license.
Both court on-prem buyers, but sell different things. Cohere sells a full retrieval stack with enterprise support. Luminal sells raw serving speed, reporting 36K tokens per second on GPT-OSS 120B across 8 H100s. A team that needs embeddings and reranking should start with Cohere; one that already has a model and needs it faster fits Luminal.
What Cohere and Luminal do
Cohere
Cohere is a Toronto-based lab that sells models and platforms to banks, governments and large enterprises rather than consumers. Its generative line is the Command family. Command A+, released May 20, 2026, is a 218B-parameter mixture-of-experts model with 25B active, published under Apache 2.0 with a 128K context window, and it combines reasoning, vision, translation and tool use in one set of weights. Command A has a 256K window and lists at $2.50 in and $10 out per million tokens, while Command R7B costs $0.0375 in. June 2026 added North Mini Code, a 30B Apache 2.0 coding model, and the lineup also includes Aya multilingual models and Transcribe for speech.
Example models: Command A+, Command A, Embed 4, Rerank 4
Full Cohere profileLuminal
Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.
Example models: GPT-OSS 120B, Llama 3 8B
Full Luminal profileShould you choose Cohere or Luminal?
Cohere
Choose Cohere for
- RAG with first-party embed and rerank models
- Private fine-tuning
- Mature on-prem and VPC deployments
Luminal
Choose Luminal for
- Maximum throughput per GPU on a self-chosen model
- Serving custom or fine-tuned architectures off any catalog
- An open-source engine teams can run on their own hardware
Cohere vs Luminal at a glance
| Attribute | ||
|---|---|---|
| Model access | Closed, plus open Command A+ | Bring your own weights |
| Flagship models | Command A+, Command A, Embed 4, Rerank 4 | No public catalog |
| Speed | 375 tok/s on Command A+ W4A4, per Cohere | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | $0.0375–$2.50 in, $0.15–$10 out per 1M | Pay per use; rates not published |
| Customization | Enterprise fine-tuning, incl. private | Compiles any PyTorch or HF model |
| Deployment | API, Bedrock, Azure, OCI, VPC, on-prem | Serverless (early access), on-prem license |
| Long context | 256K on Command A; 128K on A+ | Unknown |
Frequently asked questions
What is the difference between Cohere and Luminal?
Cohere sells enterprise models for RAG and search that run in a VPC or on-prem. Luminal sells a compiler that speeds up models you choose.
When should I choose Cohere over Luminal?
RAG with first-party embed and rerank models; Private fine-tuning; Mature on-prem and VPC deployments.
When should I choose Luminal over Cohere?
Maximum throughput per GPU on a self-chosen model; Serving custom or fine-tuned architectures off any catalog; An open-source engine teams can run on their own hardware.
Is Cohere or Luminal cheaper?
Cohere: $0.0375–$2.50 in, $0.15–$10 out per 1M. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.