Thinking Machines vs Luminal
Thinking Machines sells Tinker for post-training open models plus its Inkling models. Luminal compiles models into faster GPU code for serving.
By The Subconscious Team · Updated
Thinking Machines vs Luminal: key differences
Thinking Machines' Tinker is an API for LoRA SFT and RL on open-weight models, and its Inkling models run on a beta serverless tier. Checkpoint sampling is not meant for user-facing traffic. Luminal sits on the serving side: its compiler turns a model into native kernels ahead of time for serverless or on-prem inference.
These fit together in sequence. Train a LoRA with Tinker, merge it, then serve it through an engine like Luminal, which reports 36K tokens per second on GPT-OSS 120B across 8 H100s. Choose Thinking Machines for post-training and Luminal for production serving.
What Thinking Machines and Luminal do
Thinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileLuminal
Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.
Example models: GPT-OSS 120B, Llama 3 8B
Full Luminal profileShould you choose Thinking Machines or Luminal?
Thinking Machines
Choose Thinking Machines for
- LoRA SFT and RL through an API
- The open Inkling models
- Writing custom training loops without managing GPUs
Luminal
Choose Luminal for
- Serving a post-trained model in production
- Maximum throughput per GPU on a self-chosen model
- On-prem deployments with custom kernel work and SLAs
Thinking Machines vs Luminal at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Bring your own weights |
| Flagship models | Inkling, Inkling-Small | No public catalog |
| Speed | Unknown | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | Pay per use; rates not published |
| Customization | LoRA SFT and RL via Tinker | Compiles any PyTorch or HF model |
| Deployment | Training API, beta serverless (Inkling only) | Serverless (early access), on-prem license |
| Long context | Inkling up to 1M; Tinker 32K–256K | Unknown |
Frequently asked questions
What is the difference between Thinking Machines and Luminal?
Thinking Machines sells Tinker for post-training open models plus its Inkling models. Luminal compiles models into faster GPU code for serving.
When should I choose Thinking Machines over Luminal?
LoRA SFT and RL through an API; The open Inkling models; Writing custom training loops without managing GPUs.
When should I choose Luminal over Thinking Machines?
Serving a post-trained model in production; Maximum throughput per GPU on a self-chosen model; On-prem deployments with custom kernel work and SLAs.
Is Thinking Machines or Luminal cheaper?
Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.
Related comparisons
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Subconscious vs Luminal
OpenAI vs Luminal
Anthropic vs Luminal
Google Vertex AI vs Luminal
Amazon Bedrock vs Luminal
Together AI vs Luminal
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.