Long-running agents deserve better inference.
vs

Wafer vs Luminal

Both are 2025 YC startups making open models faster. Wafer tunes stock engines with agents; Luminal replaces the engine with a compiler.

By The Subconscious Team · Updated

Wafer vs Luminal: key differences

Wafer and Luminal came out of the same Y Combinator batch with the same target, the overhead in stock vLLM and SGLang, and different methods. Wafer's agents tune batching, decoding, quantization and kernels around a workload and keep re-tuning, reporting Qwen 3.5 397B 2.8x faster than stock SGLang. Luminal compiles the model itself: it lowers to 15 primitive ops, searches for fused kernels and emits native code ahead of time, reporting GPT-OSS 120B at 36K tokens per second on 8 H100s versus 26K for vLLM.

The products differ more than the pitch. Wafer sells Wafer Pass, flat-rate access from $10 a week that plugs into Claude Code and Cline, plus dedicated deployments on NVIDIA or AMD. Luminal sells early-access serverless endpoints for models you bring and an on-prem license, and its compiler is open source. Both speedups are self-reported against stock baselines.

What Wafer and Luminal do

Wafer

Wafer builds AI agents that act as GPU performance engineers, then sells inference on the stacks those agents tune. The company came out of Y Combinator's Summer 2025 batch as a "Cursor for CUDA" that turned slow PyTorch into custom kernels. Founders Emilio Andere and Steven Arellano are based in San Francisco. Its agents profile a workload, generate candidate configs across batching, decoding, quantization, engines, kernels and hardware, measure each one and deploy the winner.

Example models: Qwen 3.5 397B Turbo, GLM 5.1 Turbo

Full Wafer profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Wafer or Luminal?

Wafer

Choose Wafer for

  • Flat-rate open models inside coding agents
  • Dedicated endpoints re-tuned to a latency SLO
  • Hedging across NVIDIA and AMD

Luminal

Choose Luminal for

  • An open-source engine teams can run on their own hardware
  • Serving custom or fine-tuned architectures off any catalog
  • On-prem deployments with custom kernel work and SLAs

Wafer vs Luminal at a glance

AttributeWaferLuminal
Model accessOpen weightsBring your own weights
Flagship modelsQwen 3.5 397B Turbo, GLM 5.1 TurboNo public catalog
Speed2–2.8x vs stock vLLM or SGLang36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
PriceWafer Pass from $10 a weekPay per use; rates not published
CustomizationAgent-tuned dedicated deploymentsCompiles any PyTorch or HF model
DeploymentServerless pass, dedicatedServerless (early access), on-prem license
Long contextVaries by modelUnknown

Frequently asked questions

What is the difference between Wafer and Luminal?

Both are 2025 YC startups making open models faster. Wafer tunes stock engines with agents; Luminal replaces the engine with a compiler.

When should I choose Wafer over Luminal?

Flat-rate open models inside coding agents; Dedicated endpoints re-tuned to a latency SLO; Hedging across NVIDIA and AMD.

When should I choose Luminal over Wafer?

An open-source engine teams can run on their own hardware; Serving custom or fine-tuned architectures off any catalog; On-prem deployments with custom kernel work and SLAs.

Is Wafer or Luminal cheaper?

Wafer: Wafer Pass from $10 a week. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.