# Thinking Machines

> Tinker, a post-training API for open-weight models, plus the open Inkling models.

Canonical: https://www.subconscious.dev/providers/thinking-machines · By The Subconscious Team · Updated September 30, 2026

- Founded: 2025
- Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
- Website: https://thinkingmachines.ai

## Overview

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

In July 2026 the lab released its own open-weight models under Apache 2.0: Inkling, a 975B-parameter mixture-of-experts model with 41B active, and Inkling-Small at 276B total and 12B active. Both accept text, image and audio input with up to 1M tokens of context. Tinker also trains Qwen3.5, Nemotron 3, GLM-5.3, Kimi K2.6, DeepSeek-V3.1 and gpt-oss. Billing is per million tokens across prefill, sample and train meters; GPT-OSS-20B runs $0.18, $0.45 and $0.40, and cached prefill is 80% off. A beta serverless API serves only the two Inkling models, with Inkling at $1.00 in and $4.05 out. An OpenAI-compatible endpoint can sample any fine-tuned checkpoint, but the docs scope it to testing and low internal traffic.

## Upsides

- Full control of the training loop, including RL, without managing GPU clusters.
- Fine-tune large MoE models like Kimi K2.6 and Inkling that are hard to train in-house.
- Own Apache 2.0 Inkling models with native image and audio input and 1M context.
- Sample from checkpoints mid-training through an OpenAI-compatible endpoint.

## Core use cases

- Research teams running custom SFT or RL post-training on open models.
- Building a task-specialized model on top of an open-weight base.
- Evaluating the Inkling models through the beta serverless API.

## Downsides

- Not a general-purpose inference host: serverless covers only Inkling models, and checkpoint sampling is not meant for user-facing traffic.
- LoRA only, and developers write their own training code rather than using a no-code fine-tuning flow.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open weights |
| Flagship models | Inkling, Inkling-Small |
| Speed | - |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | LoRA SFT and RL via Tinker |
| Deployment | Training API, beta serverless (Inkling only) |
| Long context | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is Thinking Machines?

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

### What is Thinking Machines best for?

Research teams running custom SFT or RL post-training on open models; Building a task-specialized model on top of an open-weight base; Evaluating the Inkling models through the beta serverless API.

### How much does Thinking Machines cost?

Thinking Machines pricing at a glance: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Rates change often, so check Thinking Machines's pricing page before committing.

### How much context does Thinking Machines support?

Thinking Machines's long-context support: Inkling up to 1M; Tinker 32K–256K.

### What are the downsides of Thinking Machines?

Not a general-purpose inference host: serverless covers only Inkling models, and checkpoint sampling is not meant for user-facing traffic; LoRA only, and developers write their own training code rather than using a no-code fine-tuning flow.

### What are the best alternatives to Thinking Machines?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with Thinking Machines on this site.

## Comparisons

- [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md)
- [OpenAI vs Thinking Machines](https://www.subconscious.dev/compare/openai-vs-thinking-machines.md)
- [Anthropic vs Thinking Machines](https://www.subconscious.dev/compare/anthropic-vs-thinking-machines.md)
- [Google Vertex AI vs Thinking Machines](https://www.subconscious.dev/compare/google-vertex-vs-thinking-machines.md)
- [Amazon Bedrock vs Thinking Machines](https://www.subconscious.dev/compare/aws-bedrock-vs-thinking-machines.md)
- [Together AI vs Thinking Machines](https://www.subconscious.dev/compare/together-ai-vs-thinking-machines.md)
- [Fireworks AI vs Thinking Machines](https://www.subconscious.dev/compare/fireworks-vs-thinking-machines.md)
- [Baseten vs Thinking Machines](https://www.subconscious.dev/compare/baseten-vs-thinking-machines.md)
- [Groq vs Thinking Machines](https://www.subconscious.dev/compare/groq-vs-thinking-machines.md)
- [Cerebras vs Thinking Machines](https://www.subconscious.dev/compare/cerebras-vs-thinking-machines.md)
- [DeepInfra vs Thinking Machines](https://www.subconscious.dev/compare/deepinfra-vs-thinking-machines.md)
- [Hugging Face Inference Providers vs Thinking Machines](https://www.subconscious.dev/compare/hugging-face-vs-thinking-machines.md)
- [Modal vs Thinking Machines](https://www.subconscious.dev/compare/modal-vs-thinking-machines.md)
- [Cloudflare Workers AI vs Thinking Machines](https://www.subconscious.dev/compare/cloudflare-workers-ai-vs-thinking-machines.md)
- [xAI vs Thinking Machines](https://www.subconscious.dev/compare/xai-vs-thinking-machines.md)
- [Mistral AI vs Thinking Machines](https://www.subconscious.dev/compare/mistral-ai-vs-thinking-machines.md)
- [DeepSeek vs Thinking Machines](https://www.subconscious.dev/compare/deepseek-vs-thinking-machines.md)
- [Moonshot AI vs Thinking Machines](https://www.subconscious.dev/compare/moonshot-ai-vs-thinking-machines.md)
- [Z.ai vs Thinking Machines](https://www.subconscious.dev/compare/z-ai-vs-thinking-machines.md)
- [Alibaba Cloud vs Thinking Machines](https://www.subconscious.dev/compare/alibaba-cloud-vs-thinking-machines.md)
- [Meta vs Thinking Machines](https://www.subconscious.dev/compare/meta-vs-thinking-machines.md)
- [Cohere vs Thinking Machines](https://www.subconscious.dev/compare/cohere-vs-thinking-machines.md)
- [SambaNova vs Thinking Machines](https://www.subconscious.dev/compare/sambanova-vs-thinking-machines.md)
- [Nebius vs Thinking Machines](https://www.subconscious.dev/compare/nebius-vs-thinking-machines.md)
- [Crusoe vs Thinking Machines](https://www.subconscious.dev/compare/crusoe-vs-thinking-machines.md)
- [fal vs Thinking Machines](https://www.subconscious.dev/compare/fal-vs-thinking-machines.md)
- [Novita AI vs Thinking Machines](https://www.subconscious.dev/compare/novita-ai-vs-thinking-machines.md)
- [Venice vs Thinking Machines](https://www.subconscious.dev/compare/venice-vs-thinking-machines.md)
- [Parasail vs Thinking Machines](https://www.subconscious.dev/compare/parasail-vs-thinking-machines.md)
- [Inference.net vs Thinking Machines](https://www.subconscious.dev/compare/inference-net-vs-thinking-machines.md)
- [GMI Cloud vs Thinking Machines](https://www.subconscious.dev/compare/gmi-cloud-vs-thinking-machines.md)
- [Thinking Machines vs Sail Research](https://www.subconscious.dev/compare/thinking-machines-vs-sail-research.md)
- [Thinking Machines vs Morph](https://www.subconscious.dev/compare/thinking-machines-vs-morph.md)
- [Thinking Machines vs Relace](https://www.subconscious.dev/compare/thinking-machines-vs-relace.md)
- [Thinking Machines vs TypeSafe AI](https://www.subconscious.dev/compare/thinking-machines-vs-typesafe-ai.md)
- [Thinking Machines vs StepFun](https://www.subconscious.dev/compare/thinking-machines-vs-stepfun.md)
- [Thinking Machines vs Runware](https://www.subconscious.dev/compare/thinking-machines-vs-runware.md)
- [Thinking Machines vs StreamLake](https://www.subconscious.dev/compare/thinking-machines-vs-streamlake.md)
- [Thinking Machines vs Wafer](https://www.subconscious.dev/compare/thinking-machines-vs-wafer.md)
- [Thinking Machines vs RunInfra](https://www.subconscious.dev/compare/thinking-machines-vs-runinfra.md)
- [Thinking Machines vs Particle.AI](https://www.subconscious.dev/compare/thinking-machines-vs-particle-ai.md)

## Sources

- [Tinker](https://thinkingmachines.ai/tinker/)
- [Tinker models and pricing](https://tinker-docs.thinkingmachines.ai/tinker/models/)
- [Inkling model card](https://thinkingmachines.ai/model-card/inkling/)
- [Tinker OpenAI-compatible API](https://tinker-docs.thinkingmachines.ai/tinker/compatible-apis/openai/)

Pricing and model lineups change often; figures are a snapshot.
