# Thinking Machines vs Morph

> Thinking Machines is a lab training big general models and selling a post-training API. Morph ships small specialist models that apply coding-agent edits at speed.

Canonical: https://www.subconscious.dev/compare/thinking-machines-vs-morph · By The Subconscious Team · Updated September 30, 2026

## How they compare

The scale gap is wide. Thinking Machines' Inkling is a 975B MoE with 41B active and 1M context, and Tinker trains large open bases like Kimi K2.6 and GLM-5.3 with custom SFT or RL loops. Morph goes small on purpose. Its Fast Apply is a 7B model trained only on code merging: a frontier model writes the changed lines, and Morph merges them into the full file at 10,500+ tokens per second with up to 98% accuracy, by Morph's figures. WarpGrep adds agentic repo search, Compact compresses context and Reflex classifies. Morph also offers fine-tuning and general chat endpoints.

For coding-agent builders, Morph slots into production today as a tool next to the main model, with an OpenAI-compatible API. It does not replace a general inference provider, and a 2 to 4% merge error rate means edits still need tests or linting. Thinking Machines would matter for a team training its own coding model, for example running RL on Qwen3.5 against a test suite. It will not serve that model at scale, though. The checkpoint endpoint is scoped to testing and low internal traffic, and serverless covers only Inkling at $1.00 in and $4.05 out.

## What each one does

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

### Morph

Morph builds small, very fast specialist models that sit beside a big coding model inside an agent. Its flagship is Fast Apply. The frontier model writes only the changed lines with // ... existing code ... markers, and Morph merges them into the full file at 10,500+ tokens per second with up to 98% accuracy. It is the same idea behind Cursor's instant apply, offered as an OpenAI-compatible API.

## Which is best, and when

### Choose Thinking Machines for

- RL-training a custom coding model
- Post-training large open MoE bases
- Multimodal input with 1M context via Inkling

### Choose Morph for

- Fast file edits inside coding agents
- Cutting frontier-model output tokens
- Agentic repo search and context compaction

## At a glance

| Attribute | Thinking Machines | Morph |
|---|---|---|
| Model access | Open weights | Specialist models |
| Flagship models | Inkling, Inkling-Small | morph-v3-fast, morph-v3-large |
| Speed | - | 10,500+ tok/s Fast Apply |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | ~40% fewer tokens than full rewrites |
| Customization | LoRA SFT and RL via Tinker | Fine-tuning offered |
| Deployment | Training API, beta serverless (Inkling only) | OpenAI-compatible API |
| Long context | Inkling up to 1M; Tinker 32K–256K | - |

## FAQ

### What is the difference between Thinking Machines and Morph?

Thinking Machines is a lab training big general models and selling a post-training API. Morph ships small specialist models that apply coding-agent edits at speed.

### When should I choose Thinking Machines over Morph?

RL-training a custom coding model; Post-training large open MoE bases; Multimodal input with 1M context via Inkling.

### When should I choose Morph over Thinking Machines?

Fast file edits inside coding agents; Cutting frontier-model output tokens; Agentic repo search and context compaction.

### Is Thinking Machines or Morph cheaper?

Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Morph: ~40% fewer tokens than full rewrites. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md), [Subconscious vs Morph](https://www.subconscious.dev/compare/subconscious-vs-morph.md).

Full profiles: [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md), [Morph](https://www.subconscious.dev/providers/morph.md).
