# Morph vs Luminal

> Morph sells small specialist models that apply code edits at 10,500+ tok/s. Luminal compiles general models into faster GPU code.

Canonical: https://www.subconscious.dev/compare/morph-vs-luminal · By The Subconscious Team · Updated September 30, 2026

## How they compare

Morph makes Fast Apply models that merge coding-agent edits into files at over 10,500 tokens per second, saving tokens compared with full rewrites. It complements a main model. Luminal is general infrastructure: a compiler that turns any model into native kernels ahead of time, reporting 36K tokens per second on GPT-OSS 120B across 8 H100s.

They solve different problems. Morph makes one narrow agent step faster and cheaper. Luminal makes serving a whole model faster, whatever the task. A coding product could use both, Morph for apply and a Luminal-compiled open model for reasoning.

## What each one does

### Morph

Morph builds small, very fast specialist models that sit beside a big coding model inside an agent. Its flagship is Fast Apply. The frontier model writes only the changed lines with // ... existing code ... markers, and Morph merges them into the full file at 10,500+ tokens per second with up to 98% accuracy. It is the same idea behind Cursor's instant apply, offered as an OpenAI-compatible API.

### Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

## Which is best, and when

### Choose Morph for

- Applying coding-agent edits fast
- Fewer tokens than full file rewrites
- A drop-in OpenAI-compatible apply API

### Choose Luminal for

- Serving the main open model behind an agent
- Maximum throughput per GPU on a self-chosen model
- On-prem deployments with custom kernel work and SLAs

## At a glance

| Attribute | Morph | Luminal |
|---|---|---|
| Model access | Specialist models | Bring your own weights |
| Flagship models | morph-v3-fast, morph-v3-large | No public catalog |
| Speed | 10,500+ tok/s Fast Apply | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | ~40% fewer tokens than full rewrites | Pay per use; rates not published |
| Customization | Fine-tuning offered | Compiles any PyTorch or HF model |
| Deployment | OpenAI-compatible API | Serverless (early access), on-prem license |
| Long context | - | - |

## FAQ

### What is the difference between Morph and Luminal?

Morph sells small specialist models that apply code edits at 10,500+ tok/s. Luminal compiles general models into faster GPU code.

### When should I choose Morph over Luminal?

Applying coding-agent edits fast; Fewer tokens than full file rewrites; A drop-in OpenAI-compatible apply API.

### When should I choose Luminal over Morph?

Serving the main open model behind an agent; Maximum throughput per GPU on a self-chosen model; On-prem deployments with custom kernel work and SLAs.

### Is Morph or Luminal cheaper?

Morph: ~40% fewer tokens than full rewrites. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Morph](https://www.subconscious.dev/compare/subconscious-vs-morph.md), [Subconscious vs Luminal](https://www.subconscious.dev/compare/subconscious-vs-luminal.md).

Full profiles: [Morph](https://www.subconscious.dev/providers/morph.md), [Luminal](https://www.subconscious.dev/providers/luminal.md).
