Long-running agents deserve better inference.
vs

Morph vs Luminal

Morph sells small specialist models that apply code edits at 10,500+ tok/s. Luminal compiles general models into faster GPU code.

By The Subconscious Team · Updated

Morph vs Luminal: key differences

Morph makes Fast Apply models that merge coding-agent edits into files at over 10,500 tokens per second, saving tokens compared with full rewrites. It complements a main model. Luminal is general infrastructure: a compiler that turns any model into native kernels ahead of time, reporting 36K tokens per second on GPT-OSS 120B across 8 H100s.

They solve different problems. Morph makes one narrow agent step faster and cheaper. Luminal makes serving a whole model faster, whatever the task. A coding product could use both, Morph for apply and a Luminal-compiled open model for reasoning.

What Morph and Luminal do

Morph

Morph builds small, very fast specialist models that sit beside a big coding model inside an agent. Its flagship is Fast Apply. The frontier model writes only the changed lines with // ... existing code ... markers, and Morph merges them into the full file at 10,500+ tokens per second with up to 98% accuracy. It is the same idea behind Cursor's instant apply, offered as an OpenAI-compatible API.

Example models: morph-v3-fast, morph-v3-large

Full Morph profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Morph or Luminal?

Morph

Choose Morph for

  • Applying coding-agent edits fast
  • Fewer tokens than full file rewrites
  • A drop-in OpenAI-compatible apply API

Luminal

Choose Luminal for

  • Serving the main open model behind an agent
  • Maximum throughput per GPU on a self-chosen model
  • On-prem deployments with custom kernel work and SLAs

Morph vs Luminal at a glance

AttributeMorphLuminal
Model accessSpecialist modelsBring your own weights
Flagship modelsmorph-v3-fast, morph-v3-largeNo public catalog
Speed10,500+ tok/s Fast Apply36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price~40% fewer tokens than full rewritesPay per use; rates not published
CustomizationFine-tuning offeredCompiles any PyTorch or HF model
DeploymentOpenAI-compatible APIServerless (early access), on-prem license
Long contextUnknownUnknown

Frequently asked questions

What is the difference between Morph and Luminal?

Morph sells small specialist models that apply code edits at 10,500+ tok/s. Luminal compiles general models into faster GPU code.

When should I choose Morph over Luminal?

Applying coding-agent edits fast; Fewer tokens than full file rewrites; A drop-in OpenAI-compatible apply API.

When should I choose Luminal over Morph?

Serving the main open model behind an agent; Maximum throughput per GPU on a self-chosen model; On-prem deployments with custom kernel work and SLAs.

Is Morph or Luminal cheaper?

Morph: ~40% fewer tokens than full rewrites. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.