Long-running agents deserve better inference.
vs

Thinking Machines vs Luminal

Thinking Machines sells Tinker for post-training open models plus its Inkling models. Luminal compiles models into faster GPU code for serving.

By The Subconscious Team · Updated

Thinking Machines vs Luminal: key differences

Thinking Machines' Tinker is an API for LoRA SFT and RL on open-weight models, and its Inkling models run on a beta serverless tier. Checkpoint sampling is not meant for user-facing traffic. Luminal sits on the serving side: its compiler turns a model into native kernels ahead of time for serverless or on-prem inference.

These fit together in sequence. Train a LoRA with Tinker, merge it, then serve it through an engine like Luminal, which reports 36K tokens per second on GPT-OSS 120B across 8 H100s. Choose Thinking Machines for post-training and Luminal for production serving.

What Thinking Machines and Luminal do

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose Thinking Machines or Luminal?

Thinking Machines

Choose Thinking Machines for

  • LoRA SFT and RL through an API
  • The open Inkling models
  • Writing custom training loops without managing GPUs

Luminal

Choose Luminal for

  • Serving a post-trained model in production
  • Maximum throughput per GPU on a self-chosen model
  • On-prem deployments with custom kernel work and SLAs

Thinking Machines vs Luminal at a glance

AttributeThinking MachinesLuminal
Model accessOpen weightsBring your own weights
Flagship modelsInkling, Inkling-SmallNo public catalog
SpeedUnknown36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
PricePer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 outPay per use; rates not published
CustomizationLoRA SFT and RL via TinkerCompiles any PyTorch or HF model
DeploymentTraining API, beta serverless (Inkling only)Serverless (early access), on-prem license
Long contextInkling up to 1M; Tinker 32K–256KUnknown

Frequently asked questions

What is the difference between Thinking Machines and Luminal?

Thinking Machines sells Tinker for post-training open models plus its Inkling models. Luminal compiles models into faster GPU code for serving.

When should I choose Thinking Machines over Luminal?

LoRA SFT and RL through an API; The open Inkling models; Writing custom training loops without managing GPUs.

When should I choose Luminal over Thinking Machines?

Serving a post-trained model in production; Maximum throughput per GPU on a self-chosen model; On-prem deployments with custom kernel work and SLAs.

Is Thinking Machines or Luminal cheaper?

Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.