OpenAI vs Thinking Machines
OpenAI rents closed frontier models over a full-featured API. Thinking Machines hands teams the training loop on open weights. The choice is between using a model and building one.
By The Subconscious Team · Updated
OpenAI vs Thinking Machines: key differences
OpenAI and Thinking Machines share founders in their history, since Mira Murati and John Schulman both came from OpenAI, but the products barely overlap. OpenAI serves closed models from GPT-6 Astra at $10 in and $50 out down to GPT-5.6 Luna at $0.20 in and $1.20 out, all with a 1.05M context window, and wraps them in the Responses API, hosted tools and the Agents SDK. OpenAI sells finished closed models rather than customization. Thinking Machines is built around customization. Tinker exposes forward_backward, optim_step, sample and save_state so researchers write their own SFT or RL loops on open models, including OpenAI's own gpt-oss, while the lab runs the distributed GPUs.
On serving, OpenAI is far ahead. Thinking Machines' beta serverless API covers only Inkling and Inkling-Small, and its checkpoint endpoint is meant for testing, not user-facing traffic. Inkling does compete on price and context, at $1.00 in and $4.05 out with up to 1M tokens and native image and audio input under Apache 2.0. OpenAI's weak spot is long agent loops, since prompts over 272K bill input at 2x. The practical split: OpenAI for production assistants and tool-heavy agents today, Thinking Machines for teams that want to own a specialized open model and control how it was trained.
What OpenAI and Thinking Machines do
OpenAI
OpenAI runs the most widely adopted closed-model API. Its September 2026 lineup has GPT-6 Astra at the top for computer use, coding and long agentic runs, priced at $10 in and $50 out per million tokens. Below it sits the GPT-5.6 family: Sol for hard professional work, Terra as the balanced default, and Luna for high-volume jobs at $0.20 in and $1.20 out. All of them carry a 1.05M token context window with up to 128K output.
Example models: GPT-6 Astra, GPT-5.6 Terra
Full OpenAI profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose OpenAI or Thinking Machines?
OpenAI
Choose OpenAI for
- Production assistants that need hosted web search and file search
- Frontier computer use and coding with GPT-6 Astra
- Cheap high-volume classification on Luna
Thinking Machines
Choose Thinking Machines for
- RL post-training on gpt-oss or other open bases
- Owning Apache 2.0 weights instead of renting closed ones
- Research teams writing their own training loops
OpenAI vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Closed, plus open gpt-oss | Open weights |
| Flagship models | GPT-6 Astra, GPT-5.6 Sol, Terra, Luna | Inkling, Inkling-Small |
| Speed | Fast mode: up to 2.5x at 2x price | Unknown |
| Price | $0.20–$10 in, $1.20–$50 out per 1M | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | N/A | LoRA SFT and RL via Tinker |
| Deployment | API, Azure OpenAI, Bedrock | Training API, beta serverless (Inkling only) |
| Long context | 1.05M; 2x input past 272K | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between OpenAI and Thinking Machines?
OpenAI rents closed frontier models over a full-featured API. Thinking Machines hands teams the training loop on open weights. The choice is between using a model and building one.
When should I choose OpenAI over Thinking Machines?
Production assistants that need hosted web search and file search; Frontier computer use and coding with GPT-6 Astra; Cheap high-volume classification on Luna.
When should I choose Thinking Machines over OpenAI?
RL post-training on gpt-oss or other open bases; Owning Apache 2.0 weights instead of renting closed ones; Research teams writing their own training loops.
Is OpenAI or Thinking Machines cheaper?
OpenAI: $0.20–$10 in, $1.20–$50 out per 1M. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, OpenAI or Thinking Machines?
OpenAI: 1.05M; 2x input past 272K. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs OpenAI
OpenAI vs Anthropic
OpenAI vs Google Vertex AI
OpenAI vs Amazon Bedrock
OpenAI vs Together AI
OpenAI vs Fireworks AI
Subconscious vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Fireworks AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.