Together AI vs Thinking Machines
Both train open models, but Together also serves them at scale, while Thinking Machines focuses on a low-level training API and its own Inkling weights.
By The Subconscious Team · Updated
Together AI vs Thinking Machines: key differences
This is the closest training overlap among the providers compared here. Together offers LoRA and full-parameter SFT from $0.48 per million training tokens, with RL in closed beta and checkpoints that deploy straight to inference. Thinking Machines' Tinker is LoRA only, but it hands developers the loop itself through four calls, forward_backward, optim_step, sample and save_state, so custom RL is the default rather than a beta feature. Tinker bills by prefill, sample and train tokens, with GPT-OSS-20B at $0.18, $0.45 and $0.40 and cached prefill 80% off. Both train large open bases; Tinker's list includes Kimi K2.6, GLM-5.3, Qwen3.5 and DeepSeek-V3.1, and Together's catalog runs past thirty open models.
Serving decides most real decisions. Together runs serverless, batch, provisioned throughput with a 99% SLA, dedicated deployments and GPU clusters from $3.19 an hour for an H100 reserved, and it added canary rollouts and A/B routing in July 2026. Thinking Machines' serverless API is beta and covers only Inkling, and its checkpoint endpoint is not meant for user-facing traffic. Its distinct asset is Inkling itself: Apache 2.0, 975B parameters with 41B active, native image and audio input and up to 1M context. Teams that want train and serve on one bill pick Together; teams that want to write their own RL loop pick Tinker.
What Together AI and Thinking Machines do
Together AI
Together AI is the broadest open-model platform in the category. One bill covers per-token serverless inference, batch at up to 50% off, provisioned throughput with a 99% SLA, dedicated deployments, raw GPU clusters, managed fine-tuning and code sandboxes for agents. The text catalog runs past thirty open models, including DeepSeek V4, Kimi K3, GLM 5.2, Qwen 3.8 and MiniMax M3, plus image, video, speech and embedding models. Token prices sit at parity with Fireworks and Baseten.
Example models: Kimi K3, DeepSeek V4 Pro
Full Together AI profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Together AI or Thinking Machines?
Together AI
Choose Together AI for
- Fine-tuning and serving a checkpoint on one platform
- Full-parameter SFT, not just LoRA
- Reserved GPU clusters for large experiments
Thinking Machines
Choose Thinking Machines for
- Hand-written RL loops as a first-class feature
- LoRA on huge MoE bases without owning GPUs
- Evaluating the open Inkling models
Together AI vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | Kimi K3, DeepSeek V4, GLM 5.2, Qwen 3.8 | Inkling, Inkling-Small |
| Speed | 0.99s TTFT on DeepSeek V4 Pro | Unknown |
| Price | Parity with Fireworks and Baseten | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | LoRA and full SFT; RL in beta | LoRA SFT and RL via Tinker |
| Deployment | Serverless, dedicated, GPU clusters | Training API, beta serverless (Inkling only) |
| Long context | 512K on DeepSeek V4 Pro | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Together AI and Thinking Machines?
Both train open models, but Together also serves them at scale, while Thinking Machines focuses on a low-level training API and its own Inkling weights.
When should I choose Together AI over Thinking Machines?
Fine-tuning and serving a checkpoint on one platform; Full-parameter SFT, not just LoRA; Reserved GPU clusters for large experiments.
When should I choose Thinking Machines over Together AI?
Hand-written RL loops as a first-class feature; LoRA on huge MoE bases without owning GPUs; Evaluating the open Inkling models.
Is Together AI or Thinking Machines cheaper?
Together AI: Parity with Fireworks and Baseten. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Together AI or Thinking Machines?
Together AI: 512K on DeepSeek V4 Pro. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Together AI
OpenAI vs Together AI
Anthropic vs Together AI
Google Vertex AI vs Together AI
Amazon Bedrock vs Together AI
Together AI vs Fireworks AI
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Fireworks AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.