# Modal vs Thinking Machines

> Modal rents GPUs by the second for any Python code, training included. Thinking Machines abstracts the GPUs away behind a four-call training API.

Canonical: https://www.subconscious.dev/compare/modal-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Both let a team post-train an open model without owning hardware, but they sit at different levels of abstraction. Modal is serverless GPU compute: decorate a Python function with gpu="H100" and it builds, schedules and autoscales the container, billing per second from $0.59 an hour on a T4 to $3.95 on an H100 at list. It runs any training code on up to 8 GPUs per container, so full-parameter fine-tuning is possible if the model fits. Thinking Machines' Tinker hides the cluster entirely. Developers call forward_backward, optim_step, sample and save_state, and the lab handles distributed training across large MoE models like Kimi K2.6 and Inkling, using LoRA adapters only.

Billing shapes the choice. Modal bills GPU time, and non-preemptible US production runs about 3.75x list, near $14.81 an hour for an H100. Tinker bills per million tokens across prefill, sample and train, which is easier to forecast from dataset size. For serving, Modal covers dedicated serverless endpoints for private fine-tunes, though keeping them warm adds cost. Thinking Machines' serverless API is beta and serves only Inkling. Modal also gives $30 of free credits a month. Teams that want total flexibility, or need to serve the result, lean Modal. Teams training models too large to shard themselves lean Tinker.

## What each one does

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Modal for

- Full-parameter training or any custom Python job
- Serving private fine-tunes on serverless endpoints
- Bursty GPU work like embeddings and transcription

### Choose Thinking Machines for

- LoRA on MoE models too large to shard in-house
- Token-based training bills instead of GPU hours
- SFT or RL loops without writing distributed code

## At a glance

| Attribute | Modal | Thinking Machines |
|---|---|---|
| Model access | Bring your own weights | Open weights |
| Flagship models | None hosted | Inkling, Inkling-Small |
| Speed | ~1s container boot | - |
| Price | Per second; H100 $3.95/hr list | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Run any training code | LoRA SFT and RL via Tinker |
| Deployment | Serverless GPU containers | Training API, beta serverless (Inkling only) |
| Long context | Depends on the model you deploy | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Modal and Thinking Machines?

Modal rents GPUs by the second for any Python code, training included. Thinking Machines abstracts the GPUs away behind a four-call training API.

### When should I choose Modal over Thinking Machines?

Full-parameter training or any custom Python job; Serving private fine-tunes on serverless endpoints; Bursty GPU work like embeddings and transcription.

### When should I choose Thinking Machines over Modal?

LoRA on MoE models too large to shard in-house; Token-based training bills instead of GPU hours; SFT or RL loops without writing distributed code.

### Is Modal or Thinking Machines cheaper?

Modal: Per second; H100 $3.95/hr list. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Modal or Thinking Machines?

Modal: Depends on the model you deploy. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Modal](https://www.subconscious.dev/providers/modal.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
