We raised $5.1M for long-running agents.
vs

Together AI vs Thinking Machines

Both train open models, but Together also serves them at scale, while Thinking Machines focuses on a low-level training API and its own Inkling weights.

By The Subconscious Team · Updated

Together AI vs Thinking Machines: key differences

This is the closest training overlap among the providers compared here. Together offers LoRA and full-parameter SFT from $0.48 per million training tokens, with RL in closed beta and checkpoints that deploy straight to inference. Thinking Machines' Tinker is LoRA only, but it hands developers the loop itself through four calls, forward_backward, optim_step, sample and save_state, so custom RL is the default rather than a beta feature. Tinker bills by prefill, sample and train tokens, with GPT-OSS-20B at $0.18, $0.45 and $0.40 and cached prefill 80% off. Both train large open bases; Tinker's list includes Kimi K2.6, GLM-5.3, Qwen3.5 and DeepSeek-V3.1, and Together's catalog runs past thirty open models.

Serving decides most real decisions. Together runs serverless, batch, provisioned throughput with a 99% SLA, dedicated deployments and GPU clusters from $3.19 an hour for an H100 reserved, and it added canary rollouts and A/B routing in July 2026. Thinking Machines' serverless API is beta and covers only Inkling, and its checkpoint endpoint is not meant for user-facing traffic. Its distinct asset is Inkling itself: Apache 2.0, 975B parameters with 41B active, native image and audio input and up to 1M context. Teams that want train and serve on one bill pick Together; teams that want to write their own RL loop pick Tinker.

What Together AI and Thinking Machines do

Together AI

Together AI is the broadest open-model platform in the category. One bill covers per-token serverless inference, batch at up to 50% off, provisioned throughput with a 99% SLA, dedicated deployments, raw GPU clusters, managed fine-tuning and code sandboxes for agents. The text catalog runs past thirty open models, including DeepSeek V4, Kimi K3, GLM 5.2, Qwen 3.8 and MiniMax M3, plus image, video, speech and embedding models. Token prices sit at parity with Fireworks and Baseten.

Example models: Kimi K3, DeepSeek V4 Pro

Full Together AI profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Together AI or Thinking Machines?

Together AI

Choose Together AI for

  • Fine-tuning and serving a checkpoint on one platform
  • Full-parameter SFT, not just LoRA
  • Reserved GPU clusters for large experiments

Thinking Machines

Choose Thinking Machines for

  • Hand-written RL loops as a first-class feature
  • LoRA on huge MoE bases without owning GPUs
  • Evaluating the open Inkling models

Together AI vs Thinking Machines at a glance

AttributeTogether AIThinking Machines
Model accessOpen weightsOpen weights
Flagship modelsKimi K3, DeepSeek V4, GLM 5.2, Qwen 3.8Inkling, Inkling-Small
Speed0.99s TTFT on DeepSeek V4 ProUnknown
PriceParity with Fireworks and BasetenPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationLoRA and full SFT; RL in betaLoRA SFT and RL via Tinker
DeploymentServerless, dedicated, GPU clustersTraining API, beta serverless (Inkling only)
Long context512K on DeepSeek V4 ProInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Together AI and Thinking Machines?

Both train open models, but Together also serves them at scale, while Thinking Machines focuses on a low-level training API and its own Inkling weights.

When should I choose Together AI over Thinking Machines?

Fine-tuning and serving a checkpoint on one platform; Full-parameter SFT, not just LoRA; Reserved GPU clusters for large experiments.

When should I choose Thinking Machines over Together AI?

Hand-written RL loops as a first-class feature; LoRA on huge MoE bases without owning GPUs; Evaluating the open Inkling models.

Is Together AI or Thinking Machines cheaper?

Together AI: Parity with Fireworks and Baseten. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Together AI or Thinking Machines?

Together AI: 512K on DeepSeek V4 Pro. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.