Long-running agents deserve better inference.
vs

Subconscious vs Luminal

Both replace stock vLLM-style serving. Luminal compiles models into faster kernels; Subconscious redesigns the runtime around agent traces past 200K tokens.

By The Subconscious Team · Updated

Subconscious vs Luminal: key differences

Luminal and Subconscious both start from the view that stock runtime engines waste hardware. Luminal attacks the per-step cost: its compiler lowers a model to 15 primitive ops, searches for fused kernels, and emits native GPU code ahead of time. It reports GPT-OSS 120B at 36K tokens per second on 8 H100s, against 26K for vLLM. Subconscious attacks what grows across steps. Its runtime prunes the KV cache and preserves suffix state, so a long agent trace stops rereading its whole context, and it delivers 2x faster task completion and 50% to 80% lower cost than standard inference, with gains that grow past 200K tokens.

The products sit at different layers. Luminal is mainly an engine: an open-source compiler, early-access serverless endpoints for models you bring, and an on-prem license. Subconscious is a managed API serving GLM 5.3 and DeepSeek V4.1 Flash, with Marathon post-trained variants, billing on processed tokens, and dedicated or on-prem deployments that run nearly any open model. Luminal's speedup is aggregate throughput on a fixed benchmark; it says little about multi-million-token agent sessions, which is where Subconscious earns its keep.

What Subconscious and Luminal do

Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

Example models: GLM 5.3, DeepSeek V4.1 Flash

Full Subconscious profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Subconscious or Luminal?

Subconscious

Choose Subconscious for

  • Coding and research agents whose traces run past 200K tokens
  • A managed API that drops into Claude Code, Codex and Cursor
  • Billing on processed tokens rather than tokens sent

Luminal

Choose Luminal for

  • Compiling a custom model into fast native GPU code
  • An open-source engine teams can run on their own hardware
  • Short, high-volume requests where raw throughput matters most

Subconscious vs Luminal at a glance

AttributeSubconsciousLuminal
Model accessOpen weightsBring your own weights
Flagship modelsGLM 5.3, DeepSeek V4.1 FlashNo public catalog
Speed2x faster task completion36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price50–80% lower cost; billed on processed tokensPay per use; rates not published
CustomizationMarathon post-trained variantsCompiles any PyTorch or HF model
DeploymentManaged API, dedicated, on-premServerless (early access), on-prem license
Long context5M+ effective contextUnknown

Frequently asked questions

What is the difference between Subconscious and Luminal?

Both replace stock vLLM-style serving. Luminal compiles models into faster kernels; Subconscious redesigns the runtime around agent traces past 200K tokens.

When should I choose Subconscious over Luminal?

Coding and research agents whose traces run past 200K tokens; A managed API that drops into Claude Code, Codex and Cursor; Billing on processed tokens rather than tokens sent.

When should I choose Luminal over Subconscious?

Compiling a custom model into fast native GPU code; An open-source engine teams can run on their own hardware; Short, high-volume requests where raw throughput matters most.

Is Subconscious or Luminal cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Run your longest agent traces on Subconscious

Point the OpenAI or Anthropic SDK, or the coding agent you already use, at Subconscious. Keep Luminal for the work it does best and send the long runs to us.