We raised $5.1M for long-running agents.
vs

Mistral AI vs Thinking Machines

Both release open weights. Mistral runs a full inference API; Thinking Machines centers on Tinker, a LoRA post-training API, and 1M-context Inkling models.

By The Subconscious Team · Updated

Mistral AI vs Thinking Machines: key differences

Thinking Machines Lab, founded by Mira Murati in 2025, is primarily a training platform. Tinker exposes four low-level calls so teams write their own supervised or RL loops on open models while the lab runs the distributed GPU work, using LoRA adapters. In July 2026 it released Inkling, a 975B MoE with 41B active, and Inkling-Small, both Apache 2.0 with text, image and audio input and up to 1M context. Its beta serverless API serves only those two, with Inkling at $1.00 in and $4.05 out per million tokens. Mistral's Medium 3.5 costs $1.50 in and $7.50 out, Large 3, also Apache 2.0, costs $0.50 in and $1.50 out, and context stops at 256K.

For customization, Thinking Machines is far ahead of self-serve Mistral. Tinker trains Qwen3.5, Nemotron 3, GLM-5.3, Kimi K2.6, DeepSeek-V3.1 and gpt-oss as well as Inkling, billed per million tokens across prefill, sample and train meters. Mistral deprecated its fine-tuning API and offers Forge only to enterprises. For serving, the positions flip. Tinker's OpenAI-compatible checkpoint endpoint is scoped to testing and low internal traffic, while Mistral runs production APIs with Batch discounts, cached input savings, EU or US regions, a Priority Tier SLA and listings on Azure, Bedrock and Vertex AI. A team might train on Tinker and serve the result on another host.

What Mistral AI and Thinking Machines do

Mistral AI

Mistral AI is a Paris lab that sells its models through La Plateforme, its own API, and releases most of them as open weights. It consolidated the lineup in 2026. Mistral Medium 3.5, released April 28, is a dense 128B model that merges instruction following, reasoning and coding into one set of weights, and it replaced both Devstral 2 and the Magistral reasoning models. It costs $1.50 in and $7.50 out per million tokens and scores 77.6% on SWE-Bench Verified by Mistral's count. Mistral Small 4, a 119B mixture-of-experts model with 6.5B active, costs $0.15 in and $0.60 out. Mistral Large 3, a 675B MoE under Apache 2.0, runs $0.50 in and $1.50 out. All three carry a 256K context window.

Example models: Mistral Medium 3.5, Mistral Small 4

Full Mistral AI profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Mistral AI or Thinking Machines?

Mistral AI

Choose Mistral AI for

  • User-facing production inference
  • Regional processing and uptime SLAs
  • Low-cost volume on Small 4 or Large 3

Thinking Machines

Choose Thinking Machines for

  • Custom SFT or RL loops on open models
  • Fine-tuning large MoE models like Kimi K2.6
  • Image and audio input with 1M context on Inkling

Mistral AI vs Thinking Machines at a glance

AttributeMistral AIThinking Machines
Model accessOpen weights, plus closed CodestralOpen weights
Flagship modelsMistral Medium 3.5, Small 4, Large 3Inkling, Inkling-Small
SpeedUnknownUnknown
Price$0.15–$1.50 in, $0.60–$7.50 out per 1MPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationForge (enterprise); fine-tuning API deprecatedLoRA SFT and RL via Tinker
DeploymentAPI, Azure, Bedrock, Vertex, self-hostTraining API, beta serverless (Inkling only)
Long context256KInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Mistral AI and Thinking Machines?

Both release open weights. Mistral runs a full inference API; Thinking Machines centers on Tinker, a LoRA post-training API, and 1M-context Inkling models.

When should I choose Mistral AI over Thinking Machines?

User-facing production inference; Regional processing and uptime SLAs; Low-cost volume on Small 4 or Large 3.

When should I choose Thinking Machines over Mistral AI?

Custom SFT or RL loops on open models; Fine-tuning large MoE models like Kimi K2.6; Image and audio input with 1M context on Inkling.

Is Mistral AI or Thinking Machines cheaper?

Mistral AI: $0.15–$1.50 in, $0.60–$7.50 out per 1M. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Mistral AI or Thinking Machines?

Mistral AI: 256K. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.