# Thinking Machines vs Luminal

> Thinking Machines sells Tinker for post-training open models plus its Inkling models. Luminal compiles models into faster GPU code for serving.

Canonical: https://www.subconscious.dev/compare/thinking-machines-vs-luminal · By The Subconscious Team · Updated September 30, 2026

## How they compare

Thinking Machines' Tinker is an API for LoRA SFT and RL on open-weight models, and its Inkling models run on a beta serverless tier. Checkpoint sampling is not meant for user-facing traffic. Luminal sits on the serving side: its compiler turns a model into native kernels ahead of time for serverless or on-prem inference.

These fit together in sequence. Train a LoRA with Tinker, merge it, then serve it through an engine like Luminal, which reports 36K tokens per second on GPT-OSS 120B across 8 H100s. Choose Thinking Machines for post-training and Luminal for production serving.

## What each one does

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

### Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

## Which is best, and when

### Choose Thinking Machines for

- LoRA SFT and RL through an API
- The open Inkling models
- Writing custom training loops without managing GPUs

### Choose Luminal for

- Serving a post-trained model in production
- Maximum throughput per GPU on a self-chosen model
- On-prem deployments with custom kernel work and SLAs

## At a glance

| Attribute | Thinking Machines | Luminal |
|---|---|---|
| Model access | Open weights | Bring your own weights |
| Flagship models | Inkling, Inkling-Small | No public catalog |
| Speed | - | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | Pay per use; rates not published |
| Customization | LoRA SFT and RL via Tinker | Compiles any PyTorch or HF model |
| Deployment | Training API, beta serverless (Inkling only) | Serverless (early access), on-prem license |
| Long context | Inkling up to 1M; Tinker 32K–256K | - |

## FAQ

### What is the difference between Thinking Machines and Luminal?

Thinking Machines sells Tinker for post-training open models plus its Inkling models. Luminal compiles models into faster GPU code for serving.

### When should I choose Thinking Machines over Luminal?

LoRA SFT and RL through an API; The open Inkling models; Writing custom training loops without managing GPUs.

### When should I choose Luminal over Thinking Machines?

Serving a post-trained model in production; Maximum throughput per GPU on a self-chosen model; On-prem deployments with custom kernel work and SLAs.

### Is Thinking Machines or Luminal cheaper?

Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md), [Subconscious vs Luminal](https://www.subconscious.dev/compare/subconscious-vs-luminal.md).

Full profiles: [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md), [Luminal](https://www.subconscious.dev/providers/luminal.md).
