Thinking Machines vs Relace
Relace sells fast utility models for coding agents, hosted or self-hosted. Thinking Machines sells a training API for building your own specialized model.
By The Subconscious Team · Updated
Thinking Machines vs Relace: key differences
Both bet that specialized models beat general ones on narrow work, but they deliver it differently. Relace ships the specialized models ready-made. relace-apply-3 merges lazy edit snippets at about 10,000 tokens per second with 128K of input and output, its agentic search explores repos in parallel, and a compaction model runs at 50,000 tokens per second. Relace says Instant Apply is over 3x faster and cheaper than a full rewrite by the big model. Thinking Machines gives teams the tools to make their own. Tinker runs LoRA SFT or RL loops on Qwen3.5, gpt-oss, Kimi K2.6 and others, billed per million tokens on separate meters.
Deployment is where Relace pulls ahead for product teams. It offers a hosted API, OpenAI-compatible and REST endpoints, an OpenRouter listing and self-hosted deployment for companies that keep code in-house. Thinking Machines has no self-hosted option, and its checkpoint endpoint is meant for testing and low internal traffic. Its open Apache 2.0 Inkling weights can be self-hosted, and at 1M context they reach well past Relace's 128K cap. Relace covers coding workflows only. Thinking Machines suits research teams that want a model shaped around their own data and rewards.
What Thinking Machines and Relace do
Thinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileRelace
Relace trains small, fast models that act as tools for coding agents. Its best-known product is Instant Apply: a frontier model writes a lazy edit snippet, and relace-apply-3 merges it into the original file at about 10,000 tokens per second with 128K tokens of input and output. Relace says this runs over 3x faster and cheaper than having the big model rewrite the file. It exposes both a REST endpoint and an OpenAI-compatible one, and the model is also listed on OpenRouter.
Example models: relace-apply-3, Relace agentic search
Full Relace profileShould you choose Thinking Machines or Relace?
Thinking Machines
Choose Thinking Machines for
- Training task-specific models on your own rewards
- LoRA post-training of large MoE models
- Open Apache 2.0 weights with 1M context
Relace
Choose Relace for
- Instant apply inside coding agents
- Fast search across large codebases
- Self-hosted tools for in-house code
Thinking Machines vs Relace at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Specialist models |
| Flagship models | Inkling, Inkling-Small | relace-apply-3, agentic search |
| Speed | Unknown | ~10,000 tok/s apply |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | 3x+ cheaper than full rewrites |
| Customization | LoRA SFT and RL via Tinker | Unknown |
| Deployment | Training API, beta serverless (Inkling only) | Hosted API or self-hosted |
| Long context | Inkling up to 1M; Tinker 32K–256K | 128K max |
Frequently asked questions
What is the difference between Thinking Machines and Relace?
Relace sells fast utility models for coding agents, hosted or self-hosted. Thinking Machines sells a training API for building your own specialized model.
When should I choose Thinking Machines over Relace?
Training task-specific models on your own rewards; LoRA post-training of large MoE models; Open Apache 2.0 weights with 1M context.
When should I choose Relace over Thinking Machines?
Instant apply inside coding agents; Fast search across large codebases; Self-hosted tools for in-house code.
Is Thinking Machines or Relace cheaper?
Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Relace: 3x+ cheaper than full rewrites. The cheaper choice depends on the model and workload.
Which has more context, Thinking Machines or Relace?
Thinking Machines: Inkling up to 1M; Tinker 32K–256K. Relace: 128K max.
Related comparisons
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Subconscious vs Relace
OpenAI vs Relace
Anthropic vs Relace
Google Vertex AI vs Relace
Amazon Bedrock vs Relace
Together AI vs Relace
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.