We raised $5.1M for long-running agents.
vs

Google Vertex AI vs Thinking Machines

Vertex AI is a sprawling enterprise platform with 200+ models and custom training on TPUs. Thinking Machines is a narrow post-training API for open weights.

By The Subconscious Team · Updated

Google Vertex AI vs Thinking Machines: key differences

Both let teams train models, but at very different scopes. Vertex AI covers custom training on GPUs or TPUs, pipelines, a feature store, model registry, evaluation and deep BigQuery integration, plus Model Garden with 200+ models including Gemini 3.8, Claude and Gemma. Gemini 3.8 Flash runs $0.75 in and $3.75 out with a 1M window. Thinking Machines does one thing: Tinker, an API with four low-level calls for writing SFT or RL loops on open models, using LoRA adapters while the lab runs distributed GPUs. It trains Qwen3.5, Nemotron 3, GLM-5.3, Kimi K2.6, DeepSeek-V3.1, gpt-oss and its own Inkling models, billed per million tokens by prefill, sample and train.

Vertex wins on breadth, governance and production serving. Its agent runtime, Memory Bank and Agent Development Kit cover deployment end to end, while Thinking Machines' serverless API is in beta and serves only Inkling. The trade-off is complexity. Vertex pricing is fragmented and hard to forecast, and pipelines and registries are Vertex-native, which creates lock-in. Tinker is simpler to reason about and focused on large MoE post-training that is hard to set up in-house. Inkling adds Apache 2.0 weights with image and audio input and up to 1M context. Google Cloud shops building governed agents fit Vertex; research teams tuning open weights fit Tinker.

What Google Vertex AI and Thinking Machines do

Google Vertex AI

Vertex AI is Google Cloud's enterprise AI platform. At Google Cloud Next on April 22, 2026, Google rebranded it the Gemini Enterprise Agent Platform with an agent-first structure, though the API endpoint and most docs still say Vertex. Model Garden offers 200+ models, including Google's Gemini 3.8 family, Anthropic's Claude models and open models like Gemma, alongside Imagen, Veo and Chirp for media and speech. Google's own TPUs sit underneath much of its first-party serving.

Example models: Gemini 3.8, Claude

Full Google Vertex AI profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Google Vertex AI or Thinking Machines?

Google Vertex AI

Choose Google Vertex AI for

  • Enterprises on Google Cloud building governed agents
  • Training close to BigQuery data on GPUs or TPUs
  • Gemini and Claude behind one platform

Thinking Machines

Choose Thinking Machines for

  • LoRA post-training on large open MoE models
  • Simple per-token training bills without MLOps setup
  • Portable Apache 2.0 weights with no cloud lock-in

Google Vertex AI vs Thinking Machines at a glance

AttributeGoogle Vertex AIThinking Machines
Model accessClosed and open, 200+ modelsOpen weights
Flagship modelsGemini 3.8 Flash, Claude, GemmaInkling, Inkling-Small
SpeedFlash tier built for low latencyUnknown
PriceGemini 3.8 Flash $0.75 in, $3.75 outPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationCustom training on GPUs or TPUsLoRA SFT and RL via Tinker
DeploymentManaged on Google CloudTraining API, beta serverless (Inkling only)
Long context1M on Gemini 3.8 FlashInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Google Vertex AI and Thinking Machines?

Vertex AI is a sprawling enterprise platform with 200+ models and custom training on TPUs. Thinking Machines is a narrow post-training API for open weights.

When should I choose Google Vertex AI over Thinking Machines?

Enterprises on Google Cloud building governed agents; Training close to BigQuery data on GPUs or TPUs; Gemini and Claude behind one platform.

When should I choose Thinking Machines over Google Vertex AI?

LoRA post-training on large open MoE models; Simple per-token training bills without MLOps setup; Portable Apache 2.0 weights with no cloud lock-in.

Is Google Vertex AI or Thinking Machines cheaper?

Google Vertex AI: Gemini 3.8 Flash $0.75 in, $3.75 out. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Google Vertex AI or Thinking Machines?

Google Vertex AI: 1M on Gemini 3.8 Flash. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.