Google Vertex AI vs Thinking Machines
Vertex AI is a sprawling enterprise platform with 200+ models and custom training on TPUs. Thinking Machines is a narrow post-training API for open weights.
By The Subconscious Team · Updated
Google Vertex AI vs Thinking Machines: key differences
Both let teams train models, but at very different scopes. Vertex AI covers custom training on GPUs or TPUs, pipelines, a feature store, model registry, evaluation and deep BigQuery integration, plus Model Garden with 200+ models including Gemini 3.8, Claude and Gemma. Gemini 3.8 Flash runs $0.75 in and $3.75 out with a 1M window. Thinking Machines does one thing: Tinker, an API with four low-level calls for writing SFT or RL loops on open models, using LoRA adapters while the lab runs distributed GPUs. It trains Qwen3.5, Nemotron 3, GLM-5.3, Kimi K2.6, DeepSeek-V3.1, gpt-oss and its own Inkling models, billed per million tokens by prefill, sample and train.
Vertex wins on breadth, governance and production serving. Its agent runtime, Memory Bank and Agent Development Kit cover deployment end to end, while Thinking Machines' serverless API is in beta and serves only Inkling. The trade-off is complexity. Vertex pricing is fragmented and hard to forecast, and pipelines and registries are Vertex-native, which creates lock-in. Tinker is simpler to reason about and focused on large MoE post-training that is hard to set up in-house. Inkling adds Apache 2.0 weights with image and audio input and up to 1M context. Google Cloud shops building governed agents fit Vertex; research teams tuning open weights fit Tinker.
What Google Vertex AI and Thinking Machines do
Google Vertex AI
Vertex AI is Google Cloud's enterprise AI platform. At Google Cloud Next on April 22, 2026, Google rebranded it the Gemini Enterprise Agent Platform with an agent-first structure, though the API endpoint and most docs still say Vertex. Model Garden offers 200+ models, including Google's Gemini 3.8 family, Anthropic's Claude models and open models like Gemma, alongside Imagen, Veo and Chirp for media and speech. Google's own TPUs sit underneath much of its first-party serving.
Example models: Gemini 3.8, Claude
Full Google Vertex AI profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Google Vertex AI or Thinking Machines?
Google Vertex AI
Choose Google Vertex AI for
- Enterprises on Google Cloud building governed agents
- Training close to BigQuery data on GPUs or TPUs
- Gemini and Claude behind one platform
Thinking Machines
Choose Thinking Machines for
- LoRA post-training on large open MoE models
- Simple per-token training bills without MLOps setup
- Portable Apache 2.0 weights with no cloud lock-in
Google Vertex AI vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Closed and open, 200+ models | Open weights |
| Flagship models | Gemini 3.8 Flash, Claude, Gemma | Inkling, Inkling-Small |
| Speed | Flash tier built for low latency | Unknown |
| Price | Gemini 3.8 Flash $0.75 in, $3.75 out | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Custom training on GPUs or TPUs | LoRA SFT and RL via Tinker |
| Deployment | Managed on Google Cloud | Training API, beta serverless (Inkling only) |
| Long context | 1M on Gemini 3.8 Flash | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Google Vertex AI and Thinking Machines?
Vertex AI is a sprawling enterprise platform with 200+ models and custom training on TPUs. Thinking Machines is a narrow post-training API for open weights.
When should I choose Google Vertex AI over Thinking Machines?
Enterprises on Google Cloud building governed agents; Training close to BigQuery data on GPUs or TPUs; Gemini and Claude behind one platform.
When should I choose Thinking Machines over Google Vertex AI?
LoRA post-training on large open MoE models; Simple per-token training bills without MLOps setup; Portable Apache 2.0 weights with no cloud lock-in.
Is Google Vertex AI or Thinking Machines cheaper?
Google Vertex AI: Gemini 3.8 Flash $0.75 in, $3.75 out. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Google Vertex AI or Thinking Machines?
Google Vertex AI: 1M on Gemini 3.8 Flash. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Google Vertex AI
OpenAI vs Google Vertex AI
Anthropic vs Google Vertex AI
Google Vertex AI vs Amazon Bedrock
Google Vertex AI vs Together AI
Google Vertex AI vs Fireworks AI
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Fireworks AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.