# Together AI vs Thinking Machines

> Both train open models, but Together also serves them at scale, while Thinking Machines focuses on a low-level training API and its own Inkling weights.

Canonical: https://www.subconscious.dev/compare/together-ai-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

This is the closest training overlap among the providers compared here. Together offers LoRA and full-parameter SFT from $0.48 per million training tokens, with RL in closed beta and checkpoints that deploy straight to inference. Thinking Machines' Tinker is LoRA only, but it hands developers the loop itself through four calls, forward_backward, optim_step, sample and save_state, so custom RL is the default rather than a beta feature. Tinker bills by prefill, sample and train tokens, with GPT-OSS-20B at $0.18, $0.45 and $0.40 and cached prefill 80% off. Both train large open bases; Tinker's list includes Kimi K2.6, GLM-5.3, Qwen3.5 and DeepSeek-V3.1, and Together's catalog runs past thirty open models.

Serving decides most real decisions. Together runs serverless, batch, provisioned throughput with a 99% SLA, dedicated deployments and GPU clusters from $3.19 an hour for an H100 reserved, and it added canary rollouts and A/B routing in July 2026. Thinking Machines' serverless API is beta and covers only Inkling, and its checkpoint endpoint is not meant for user-facing traffic. Its distinct asset is Inkling itself: Apache 2.0, 975B parameters with 41B active, native image and audio input and up to 1M context. Teams that want train and serve on one bill pick Together; teams that want to write their own RL loop pick Tinker.

## What each one does

### Together AI

Together AI is the broadest open-model platform in the category. One bill covers per-token serverless inference, batch at up to 50% off, provisioned throughput with a 99% SLA, dedicated deployments, raw GPU clusters, managed fine-tuning and code sandboxes for agents. The text catalog runs past thirty open models, including DeepSeek V4, Kimi K3, GLM 5.2, Qwen 3.8 and MiniMax M3, plus image, video, speech and embedding models. Token prices sit at parity with Fireworks and Baseten.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Together AI for

- Fine-tuning and serving a checkpoint on one platform
- Full-parameter SFT, not just LoRA
- Reserved GPU clusters for large experiments

### Choose Thinking Machines for

- Hand-written RL loops as a first-class feature
- LoRA on huge MoE bases without owning GPUs
- Evaluating the open Inkling models

## At a glance

| Attribute | Together AI | Thinking Machines |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | Kimi K3, DeepSeek V4, GLM 5.2, Qwen 3.8 | Inkling, Inkling-Small |
| Speed | 0.99s TTFT on DeepSeek V4 Pro | - |
| Price | Parity with Fireworks and Baseten | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | LoRA and full SFT; RL in beta | LoRA SFT and RL via Tinker |
| Deployment | Serverless, dedicated, GPU clusters | Training API, beta serverless (Inkling only) |
| Long context | 512K on DeepSeek V4 Pro | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Together AI and Thinking Machines?

Both train open models, but Together also serves them at scale, while Thinking Machines focuses on a low-level training API and its own Inkling weights.

### When should I choose Together AI over Thinking Machines?

Fine-tuning and serving a checkpoint on one platform; Full-parameter SFT, not just LoRA; Reserved GPU clusters for large experiments.

### When should I choose Thinking Machines over Together AI?

Hand-written RL loops as a first-class feature; LoRA on huge MoE bases without owning GPUs; Evaluating the open Inkling models.

### Is Together AI or Thinking Machines cheaper?

Together AI: Parity with Fireworks and Baseten. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Together AI or Thinking Machines?

Together AI: 512K on DeepSeek V4 Pro. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Together AI](https://www.subconscious.dev/compare/subconscious-vs-together-ai.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [Together AI](https://www.subconscious.dev/providers/together-ai.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
