Thinking Machines vs Morph
Thinking Machines is a lab training big general models and selling a post-training API. Morph ships small specialist models that apply coding-agent edits at speed.
By The Subconscious Team · Updated
Thinking Machines vs Morph: key differences
The scale gap is wide. Thinking Machines' Inkling is a 975B MoE with 41B active and 1M context, and Tinker trains large open bases like Kimi K2.6 and GLM-5.3 with custom SFT or RL loops. Morph goes small on purpose. Its Fast Apply is a 7B model trained only on code merging: a frontier model writes the changed lines, and Morph merges them into the full file at 10,500+ tokens per second with up to 98% accuracy, by Morph's figures. WarpGrep adds agentic repo search, Compact compresses context and Reflex classifies. Morph also offers fine-tuning and general chat endpoints.
For coding-agent builders, Morph slots into production today as a tool next to the main model, with an OpenAI-compatible API. It does not replace a general inference provider, and a 2 to 4% merge error rate means edits still need tests or linting. Thinking Machines would matter for a team training its own coding model, for example running RL on Qwen3.5 against a test suite. It will not serve that model at scale, though. The checkpoint endpoint is scoped to testing and low internal traffic, and serverless covers only Inkling at $1.00 in and $4.05 out.
What Thinking Machines and Morph do
Thinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileMorph
Morph builds small, very fast specialist models that sit beside a big coding model inside an agent. Its flagship is Fast Apply. The frontier model writes only the changed lines with // ... existing code ... markers, and Morph merges them into the full file at 10,500+ tokens per second with up to 98% accuracy. It is the same idea behind Cursor's instant apply, offered as an OpenAI-compatible API.
Example models: morph-v3-fast, morph-v3-large
Full Morph profileShould you choose Thinking Machines or Morph?
Thinking Machines
Choose Thinking Machines for
- RL-training a custom coding model
- Post-training large open MoE bases
- Multimodal input with 1M context via Inkling
Morph
Choose Morph for
- Fast file edits inside coding agents
- Cutting frontier-model output tokens
- Agentic repo search and context compaction
Thinking Machines vs Morph at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Specialist models |
| Flagship models | Inkling, Inkling-Small | morph-v3-fast, morph-v3-large |
| Speed | Unknown | 10,500+ tok/s Fast Apply |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | ~40% fewer tokens than full rewrites |
| Customization | LoRA SFT and RL via Tinker | Fine-tuning offered |
| Deployment | Training API, beta serverless (Inkling only) | OpenAI-compatible API |
| Long context | Inkling up to 1M; Tinker 32K–256K | Unknown |
Frequently asked questions
What is the difference between Thinking Machines and Morph?
Thinking Machines is a lab training big general models and selling a post-training API. Morph ships small specialist models that apply coding-agent edits at speed.
When should I choose Thinking Machines over Morph?
RL-training a custom coding model; Post-training large open MoE bases; Multimodal input with 1M context via Inkling.
When should I choose Morph over Thinking Machines?
Fast file edits inside coding agents; Cutting frontier-model output tokens; Agentic repo search and context compaction.
Is Thinking Machines or Morph cheaper?
Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Morph: ~40% fewer tokens than full rewrites. The cheaper choice depends on the model and workload.
Related comparisons
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Subconscious vs Morph
OpenAI vs Morph
Anthropic vs Morph
Google Vertex AI vs Morph
Amazon Bedrock vs Morph
Together AI vs Morph
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.