Long-running agents deserve better inference.
vs

Relace vs Luminal

Relace sells small fast models for coding-agent apply, search and compaction. Luminal compiles general models into faster GPU code.

By The Subconscious Team · Updated

Relace vs Luminal: key differences

Relace builds specialist models that act as tools for coding agents, applying edits at about 10,000 tokens per second, plus agentic search and compaction, hosted or self-hosted, with a 128K cap. Luminal is not a model company. Its compiler turns a model into fused native kernels ahead of time and sells serverless or on-prem serving.

Relace speeds up specific agent steps; Luminal speeds up the engine underneath any model, reporting 36K tokens per second on GPT-OSS 120B over 8 H100s. Teams building coding agents may use Relace for tool calls and Luminal for their main open model.

What Relace and Luminal do

Relace

Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.

Example models: relace-apply-3, Relace agentic search

Full Relace profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Relace or Luminal?

Relace

Choose Relace for

  • Fast apply, search and compaction for coding agents
  • Self-hosting specialist models
  • Cheaper edits than full rewrites

Luminal

Choose Luminal for

  • Serving the main open model behind an agent
  • On-prem deployments with custom kernel work and SLAs
  • An open-source engine teams can run on their own hardware

Relace vs Luminal at a glance

AttributeRelaceLuminal
Model accessSpecialist modelsBring your own weights
Flagship modelsrelace-apply-3, agentic searchNo public catalog
Speed~10,000 tok/s apply36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price3x+ cheaper than full rewritesPay per use; rates not published
CustomizationUnknownCompiles any PyTorch or HF model
DeploymentHosted API or self-hostedServerless (early access), on-prem license
Long context128K maxUnknown

Frequently asked questions

What is the difference between Relace and Luminal?

Relace sells small fast models for coding-agent apply, search and compaction. Luminal compiles general models into faster GPU code.

When should I choose Relace over Luminal?

Fast apply, search and compaction for coding agents; Self-hosting specialist models; Cheaper edits than full rewrites.

When should I choose Luminal over Relace?

Serving the main open model behind an agent; On-prem deployments with custom kernel work and SLAs; An open-source engine teams can run on their own hardware.

Is Relace or Luminal cheaper?

Relace: 3x+ cheaper than full rewrites. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.