# Google Vertex AI vs Thinking Machines

> Vertex AI is a sprawling enterprise platform with 200+ models and custom training on TPUs. Thinking Machines is a narrow post-training API for open weights.

Canonical: https://www.subconscious.dev/compare/google-vertex-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

Both let teams train models, but at very different scopes. Vertex AI covers custom training on GPUs or TPUs, pipelines, a feature store, model registry, evaluation and deep BigQuery integration, plus Model Garden with 200+ models including Gemini 3.8, Claude and Gemma. Gemini 3.8 Flash runs $0.75 in and $3.75 out with a 1M window. Thinking Machines does one thing: Tinker, an API with four low-level calls for writing SFT or RL loops on open models, using LoRA adapters while the lab runs distributed GPUs. It trains Qwen3.5, Nemotron 3, GLM-5.3, Kimi K2.6, DeepSeek-V3.1, gpt-oss and its own Inkling models, billed per million tokens by prefill, sample and train.

Vertex wins on breadth, governance and production serving. Its agent runtime, Memory Bank and Agent Development Kit cover deployment end to end, while Thinking Machines' serverless API is in beta and serves only Inkling. The trade-off is complexity. Vertex pricing is fragmented and hard to forecast, and pipelines and registries are Vertex-native, which creates lock-in. Tinker is simpler to reason about and focused on large MoE post-training that is hard to set up in-house. Inkling adds Apache 2.0 weights with image and audio input and up to 1M context. Google Cloud shops building governed agents fit Vertex; research teams tuning open weights fit Tinker.

## What each one does

### Google Vertex AI

Vertex AI is Google Cloud's enterprise AI platform. At Google Cloud Next on April 22, 2026, Google rebranded it the Gemini Enterprise Agent Platform with an agent-first structure, though the API endpoint and most docs still say Vertex. Model Garden offers 200+ models, including Google's Gemini 3.8 family, Anthropic's Claude models and open models like Gemma, alongside Imagen, Veo and Chirp for media and speech. Google's own TPUs sit underneath much of its first-party serving.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Google Vertex AI for

- Enterprises on Google Cloud building governed agents
- Training close to BigQuery data on GPUs or TPUs
- Gemini and Claude behind one platform

### Choose Thinking Machines for

- LoRA post-training on large open MoE models
- Simple per-token training bills without MLOps setup
- Portable Apache 2.0 weights with no cloud lock-in

## At a glance

| Attribute | Google Vertex AI | Thinking Machines |
|---|---|---|
| Model access | Closed and open, 200+ models | Open weights |
| Flagship models | Gemini 3.8 Flash, Claude, Gemma | Inkling, Inkling-Small |
| Speed | Flash tier built for low latency | - |
| Price | Gemini 3.8 Flash $0.75 in, $3.75 out | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Custom training on GPUs or TPUs | LoRA SFT and RL via Tinker |
| Deployment | Managed on Google Cloud | Training API, beta serverless (Inkling only) |
| Long context | 1M on Gemini 3.8 Flash | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Google Vertex AI and Thinking Machines?

Vertex AI is a sprawling enterprise platform with 200+ models and custom training on TPUs. Thinking Machines is a narrow post-training API for open weights.

### When should I choose Google Vertex AI over Thinking Machines?

Enterprises on Google Cloud building governed agents; Training close to BigQuery data on GPUs or TPUs; Gemini and Claude behind one platform.

### When should I choose Thinking Machines over Google Vertex AI?

LoRA post-training on large open MoE models; Simple per-token training bills without MLOps setup; Portable Apache 2.0 weights with no cloud lock-in.

### Is Google Vertex AI or Thinking Machines cheaper?

Google Vertex AI: Gemini 3.8 Flash $0.75 in, $3.75 out. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Google Vertex AI or Thinking Machines?

Google Vertex AI: 1M on Gemini 3.8 Flash. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Google Vertex AI](https://www.subconscious.dev/compare/subconscious-vs-google-vertex.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Google Vertex AI](https://www.subconscious.dev/providers/google-vertex.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
