We raised $5.1M for long-running agents.
vs

Thinking Machines vs Morph

Thinking Machines is a lab training big general models and selling a post-training API. Morph ships small specialist models that apply coding-agent edits at speed.

By The Subconscious Team · Updated

Thinking Machines vs Morph: key differences

The scale gap is wide. Thinking Machines' Inkling is a 975B MoE with 41B active and 1M context, and Tinker trains large open bases like Kimi K2.6 and GLM-5.3 with custom SFT or RL loops. Morph goes small on purpose. Its Fast Apply is a 7B model trained only on code merging: a frontier model writes the changed lines, and Morph merges them into the full file at 10,500+ tokens per second with up to 98% accuracy, by Morph's figures. WarpGrep adds agentic repo search, Compact compresses context and Reflex classifies. Morph also offers fine-tuning and general chat endpoints.

For coding-agent builders, Morph slots into production today as a tool next to the main model, with an OpenAI-compatible API. It does not replace a general inference provider, and a 2 to 4% merge error rate means edits still need tests or linting. Thinking Machines would matter for a team training its own coding model, for example running RL on Qwen3.5 against a test suite. It will not serve that model at scale, though. The checkpoint endpoint is scoped to testing and low internal traffic, and serverless covers only Inkling at $1.00 in and $4.05 out.

What Thinking Machines and Morph do

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Morph

Morph builds small, very fast specialist models that sit beside a big coding model inside an agent. Its flagship is Fast Apply. The frontier model writes only the changed lines with // ... existing code ... markers, and Morph merges them into the full file at 10,500+ tokens per second with up to 98% accuracy. It is the same idea behind Cursor's instant apply, offered as an OpenAI-compatible API.

Example models: morph-v3-fast, morph-v3-large

Full Morph profile

Should you choose Thinking Machines or Morph?

Thinking Machines

Choose Thinking Machines for

  • RL-training a custom coding model
  • Post-training large open MoE bases
  • Multimodal input with 1M context via Inkling

Morph

Choose Morph for

  • Fast file edits inside coding agents
  • Cutting frontier-model output tokens
  • Agentic repo search and context compaction

Thinking Machines vs Morph at a glance

AttributeThinking MachinesMorph
Model accessOpen weightsSpecialist models
Flagship modelsInkling, Inkling-Smallmorph-v3-fast, morph-v3-large
SpeedUnknown10,500+ tok/s Fast Apply
PricePer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out~40% fewer tokens than full rewrites
CustomizationLoRA SFT and RL via TinkerFine-tuning offered
DeploymentTraining API, beta serverless (Inkling only)OpenAI-compatible API
Long contextInkling up to 1M; Tinker 32K–256KUnknown

Frequently asked questions

What is the difference between Thinking Machines and Morph?

Thinking Machines is a lab training big general models and selling a post-training API. Morph ships small specialist models that apply coding-agent edits at speed.

When should I choose Thinking Machines over Morph?

RL-training a custom coding model; Post-training large open MoE bases; Multimodal input with 1M context via Inkling.

When should I choose Morph over Thinking Machines?

Fast file edits inside coding agents; Cutting frontier-model output tokens; Agentic repo search and context compaction.

Is Thinking Machines or Morph cheaper?

Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Morph: ~40% fewer tokens than full rewrites. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.