We raised $5.1M for long-running agents.
vs

Modal vs Thinking Machines

Modal rents GPUs by the second for any Python code, training included. Thinking Machines abstracts the GPUs away behind a four-call training API.

By The Subconscious Team · Updated

Modal vs Thinking Machines: key differences

Both let a team post-train an open model without owning hardware, but they sit at different levels of abstraction. Modal is serverless GPU compute: decorate a Python function with gpu="H100" and it builds, schedules and autoscales the container, billing per second from $0.59 an hour on a T4 to $3.95 on an H100 at list. It runs any training code on up to 8 GPUs per container, so full-parameter fine-tuning is possible if the model fits. Thinking Machines' Tinker hides the cluster entirely. Developers call forward_backward, optim_step, sample and save_state, and the lab handles distributed training across large MoE models like Kimi K2.6 and Inkling, using LoRA adapters only.

Billing shapes the choice. Modal bills GPU time, and non-preemptible US production runs about 3.75x list, near $14.81 an hour for an H100. Tinker bills per million tokens across prefill, sample and train, which is easier to forecast from dataset size. For serving, Modal covers dedicated serverless endpoints for private fine-tunes, though keeping them warm adds cost. Thinking Machines' serverless API is beta and serves only Inkling. Modal also gives $30 of free credits a month. Teams that want total flexibility, or need to serve the result, lean Modal. Teams training models too large to shard themselves lean Tinker.

What Modal and Thinking Machines do

Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper

Full Modal profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Modal or Thinking Machines?

Modal

Choose Modal for

  • Full-parameter training or any custom Python job
  • Serving private fine-tunes on serverless endpoints
  • Bursty GPU work like embeddings and transcription

Thinking Machines

Choose Thinking Machines for

  • LoRA on MoE models too large to shard in-house
  • Token-based training bills instead of GPU hours
  • SFT or RL loops without writing distributed code

Modal vs Thinking Machines at a glance

AttributeModalThinking Machines
Model accessBring your own weightsOpen weights
Flagship modelsNone hostedInkling, Inkling-Small
Speed~1s container bootUnknown
PricePer second; H100 $3.95/hr listPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationRun any training codeLoRA SFT and RL via Tinker
DeploymentServerless GPU containersTraining API, beta serverless (Inkling only)
Long contextDepends on the model you deployInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Modal and Thinking Machines?

Modal rents GPUs by the second for any Python code, training included. Thinking Machines abstracts the GPUs away behind a four-call training API.

When should I choose Modal over Thinking Machines?

Full-parameter training or any custom Python job; Serving private fine-tunes on serverless endpoints; Bursty GPU work like embeddings and transcription.

When should I choose Thinking Machines over Modal?

LoRA on MoE models too large to shard in-house; Token-based training bills instead of GPU hours; SFT or RL loops without writing distributed code.

Is Modal or Thinking Machines cheaper?

Modal: Per second; H100 $3.95/hr list. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Modal or Thinking Machines?

Modal: Depends on the model you deploy. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.