Mistral AI vs Thinking Machines
Both release open weights. Mistral runs a full inference API; Thinking Machines centers on Tinker, a LoRA post-training API, and 1M-context Inkling models.
By The Subconscious Team · Updated
Mistral AI vs Thinking Machines: key differences
Thinking Machines Lab, founded by Mira Murati in 2025, is primarily a training platform. Tinker exposes four low-level calls so teams write their own supervised or RL loops on open models while the lab runs the distributed GPU work, using LoRA adapters. In July 2026 it released Inkling, a 975B MoE with 41B active, and Inkling-Small, both Apache 2.0 with text, image and audio input and up to 1M context. Its beta serverless API serves only those two, with Inkling at $1.00 in and $4.05 out per million tokens. Mistral's Medium 3.5 costs $1.50 in and $7.50 out, Large 3, also Apache 2.0, costs $0.50 in and $1.50 out, and context stops at 256K.
For customization, Thinking Machines is far ahead of self-serve Mistral. Tinker trains Qwen3.5, Nemotron 3, GLM-5.3, Kimi K2.6, DeepSeek-V3.1 and gpt-oss as well as Inkling, billed per million tokens across prefill, sample and train meters. Mistral deprecated its fine-tuning API and offers Forge only to enterprises. For serving, the positions flip. Tinker's OpenAI-compatible checkpoint endpoint is scoped to testing and low internal traffic, while Mistral runs production APIs with Batch discounts, cached input savings, EU or US regions, a Priority Tier SLA and listings on Azure, Bedrock and Vertex AI. A team might train on Tinker and serve the result on another host.
What Mistral AI and Thinking Machines do
Mistral AI
Mistral AI is a Paris lab that sells its models through La Plateforme, its own API, and releases most of them as open weights. It consolidated the lineup in 2026. Mistral Medium 3.5, released April 28, is a dense 128B model that merges instruction following, reasoning and coding into one set of weights, and it replaced both Devstral 2 and the Magistral reasoning models. It costs $1.50 in and $7.50 out per million tokens and scores 77.6% on SWE-Bench Verified by Mistral's count. Mistral Small 4, a 119B mixture-of-experts model with 6.5B active, costs $0.15 in and $0.60 out. Mistral Large 3, a 675B MoE under Apache 2.0, runs $0.50 in and $1.50 out. All three carry a 256K context window.
Example models: Mistral Medium 3.5, Mistral Small 4
Full Mistral AI profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Mistral AI or Thinking Machines?
Mistral AI
Choose Mistral AI for
- User-facing production inference
- Regional processing and uptime SLAs
- Low-cost volume on Small 4 or Large 3
Thinking Machines
Choose Thinking Machines for
- Custom SFT or RL loops on open models
- Fine-tuning large MoE models like Kimi K2.6
- Image and audio input with 1M context on Inkling
Mistral AI vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights, plus closed Codestral | Open weights |
| Flagship models | Mistral Medium 3.5, Small 4, Large 3 | Inkling, Inkling-Small |
| Speed | Unknown | Unknown |
| Price | $0.15–$1.50 in, $0.60–$7.50 out per 1M | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Forge (enterprise); fine-tuning API deprecated | LoRA SFT and RL via Tinker |
| Deployment | API, Azure, Bedrock, Vertex, self-host | Training API, beta serverless (Inkling only) |
| Long context | 256K | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Mistral AI and Thinking Machines?
Both release open weights. Mistral runs a full inference API; Thinking Machines centers on Tinker, a LoRA post-training API, and 1M-context Inkling models.
When should I choose Mistral AI over Thinking Machines?
User-facing production inference; Regional processing and uptime SLAs; Low-cost volume on Small 4 or Large 3.
When should I choose Thinking Machines over Mistral AI?
Custom SFT or RL loops on open models; Fine-tuning large MoE models like Kimi K2.6; Image and audio input with 1M context on Inkling.
Is Mistral AI or Thinking Machines cheaper?
Mistral AI: $0.15–$1.50 in, $0.60–$7.50 out per 1M. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Mistral AI or Thinking Machines?
Mistral AI: 256K. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Mistral AI
OpenAI vs Mistral AI
Anthropic vs Mistral AI
Google Vertex AI vs Mistral AI
Amazon Bedrock vs Mistral AI
Together AI vs Mistral AI
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.