# Subconscious vs Luminal

> Both replace stock vLLM-style serving. Luminal compiles models into faster kernels; Subconscious redesigns the runtime around agent traces past 200K tokens.

Canonical: https://www.subconscious.dev/compare/subconscious-vs-luminal · By The Subconscious Team · Updated September 30, 2026

## How they compare

Luminal and Subconscious both start from the view that stock runtime engines waste hardware. Luminal attacks the per-step cost: its compiler lowers a model to 15 primitive ops, searches for fused kernels, and emits native GPU code ahead of time. It reports GPT-OSS 120B at 36K tokens per second on 8 H100s, against 26K for vLLM. Subconscious attacks what grows across steps. Its runtime prunes the KV cache and preserves suffix state, so a long agent trace stops rereading its whole context, and it delivers 2x faster task completion and 50% to 80% lower cost than standard inference, with gains that grow past 200K tokens.

The products sit at different layers. Luminal is mainly an engine: an open-source compiler, early-access serverless endpoints for models you bring, and an on-prem license. Subconscious is a managed API serving GLM 5.3 and DeepSeek V4.1 Flash, with Marathon post-trained variants, billing on processed tokens, and dedicated or on-prem deployments that run nearly any open model. Luminal's speedup is aggregate throughput on a fixed benchmark; it says little about multi-million-token agent sessions, which is where Subconscious earns its keep.

## What each one does

### Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

## Which is best, and when

### Choose Subconscious for

- Coding and research agents whose traces run past 200K tokens
- A managed API that drops into Claude Code, Codex and Cursor
- Billing on processed tokens rather than tokens sent

### Choose Luminal for

- Compiling a custom model into fast native GPU code
- An open-source engine teams can run on their own hardware
- Short, high-volume requests where raw throughput matters most

## At a glance

| Attribute | Subconscious | Luminal |
|---|---|---|
| Model access | Open weights | Bring your own weights |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | No public catalog |
| Speed | 2x faster task completion | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | 50–80% lower cost; billed on processed tokens | Pay per use; rates not published |
| Customization | Marathon post-trained variants | Compiles any PyTorch or HF model |
| Deployment | Managed API, dedicated, on-prem | Serverless (early access), on-prem license |
| Long context | 5M+ effective context | - |

## FAQ

### What is the difference between Subconscious and Luminal?

Both replace stock vLLM-style serving. Luminal compiles models into faster kernels; Subconscious redesigns the runtime around agent traces past 200K tokens.

### When should I choose Subconscious over Luminal?

Coding and research agents whose traces run past 200K tokens; A managed API that drops into Claude Code, Codex and Cursor; Billing on processed tokens rather than tokens sent.

### When should I choose Luminal over Subconscious?

Compiling a custom model into fast native GPU code; An open-source engine teams can run on their own hardware; Short, high-volume requests where raw throughput matters most.

### Is Subconscious or Luminal cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

## Try Subconscious

Subconscious speaks the OpenAI and Anthropic API formats. Base URL: https://api.subconscious.dev/v1. Docs: https://docs.subconscious.dev. Get an API key: https://platform.subconscious.dev/signin. Agent guide: https://www.subconscious.dev/agents.md.

Full profiles: [Subconscious](https://www.subconscious.dev/providers/subconscious.md), [Luminal](https://www.subconscious.dev/providers/luminal.md).
