Subconscious vs Luminal
Both replace stock vLLM-style serving. Luminal compiles models into faster kernels; Subconscious redesigns the runtime around agent traces past 200K tokens.
By The Subconscious Team · Updated
Subconscious vs Luminal: key differences
Luminal and Subconscious both start from the view that stock runtime engines waste hardware. Luminal attacks the per-step cost: its compiler lowers a model to 15 primitive ops, searches for fused kernels, and emits native GPU code ahead of time. It reports GPT-OSS 120B at 36K tokens per second on 8 H100s, against 26K for vLLM. Subconscious attacks what grows across steps. Its runtime prunes the KV cache and preserves suffix state, so a long agent trace stops rereading its whole context, and it delivers 2x faster task completion and 50% to 80% lower cost than standard inference, with gains that grow past 200K tokens.
The products sit at different layers. Luminal is mainly an engine: an open-source compiler, early-access serverless endpoints for models you bring, and an on-prem license. Subconscious is a managed API serving GLM 5.3 and DeepSeek V4.1 Flash, with Marathon post-trained variants, billing on processed tokens, and dedicated or on-prem deployments that run nearly any open model. Luminal's speedup is aggregate throughput on a fixed benchmark; it says little about multi-million-token agent sessions, which is where Subconscious earns its keep.
What Subconscious and Luminal do
Subconscious
Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.
Example models: GLM 5.3, DeepSeek V4.1 Flash
Full Subconscious profileLuminal
Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.
Example models: GPT-OSS 120B, Llama 3 8B
Full Luminal profileShould you choose Subconscious or Luminal?
Subconscious
Choose Subconscious for
- Coding and research agents whose traces run past 200K tokens
- A managed API that drops into Claude Code, Codex and Cursor
- Billing on processed tokens rather than tokens sent
Luminal
Choose Luminal for
- Compiling a custom model into fast native GPU code
- An open-source engine teams can run on their own hardware
- Short, high-volume requests where raw throughput matters most
Subconscious vs Luminal at a glance
| Attribute | ||
|---|---|---|
| Model access | Open weights | Bring your own weights |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | No public catalog |
| Speed | 2x faster task completion | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | 50–80% lower cost; billed on processed tokens | Pay per use; rates not published |
| Customization | Marathon post-trained variants | Compiles any PyTorch or HF model |
| Deployment | Managed API, dedicated, on-prem | Serverless (early access), on-prem license |
| Long context | 5M+ effective context | Unknown |
Frequently asked questions
What is the difference between Subconscious and Luminal?
Both replace stock vLLM-style serving. Luminal compiles models into faster kernels; Subconscious redesigns the runtime around agent traces past 200K tokens.
When should I choose Subconscious over Luminal?
Coding and research agents whose traces run past 200K tokens; A managed API that drops into Claude Code, Codex and Cursor; Billing on processed tokens rather than tokens sent.
When should I choose Luminal over Subconscious?
Compiling a custom model into fast native GPU code; An open-source engine teams can run on their own hardware; Short, high-volume requests where raw throughput matters most.
Is Subconscious or Luminal cheaper?
Subconscious: 50–80% lower cost; billed on processed tokens. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.
Related comparisons
Run your longest agent traces on Subconscious
Point the OpenAI or Anthropic SDK, or the coding agent you already use, at Subconscious. Keep Luminal for the work it does best and send the long runs to us.