SambaNova vs Thinking Machines
SambaNova serves big open models fast on its own dataflow chip. Thinking Machines is mainly a place to train them, with Tinker for LoRA post-training and a beta API for Inkling.
By The Subconscious Team · Updated
SambaNova vs Thinking Machines: key differences
These two sit at opposite ends of the model lifecycle. SambaNova's SambaCloud serves MiniMax M2.7, DeepSeek, Gemma 4 31B and GPT-OSS 120B on its RDU chip, with GPT-OSS 120B at $0.22 in and $0.59 out. SambaNova claims its SN50 rack runs MiniMax M2.7 near 820 tokens per second in its fastest configuration, and its three-tier memory lets one system hot swap models in milliseconds. It offers no public fine-tuning. Thinking Machines is built around Tinker, where teams write their own SFT or RL loops with four low-level calls and the lab runs the distributed GPU work on LoRA adapters. Its only production-style serving is a beta serverless API for Inkling and Inkling-Small.
Context and cost differ too. Inkling accepts text, image and audio with up to 1M tokens, while SambaCloud tops out around 192K on MiniMax M2.7. Inkling costs $1.00 in and $4.05 out, well above SambaNova's GPT-OSS rates. Tinker bills separate prefill, sample and train meters, so the spend tracks experiments rather than traffic. Its OpenAI-compatible checkpoint endpoint is documented for testing and low internal traffic only. Teams that need a custom model can train on Tinker, but they will need another host for fast user-facing serving, and SambaNova does not advertise hosting customer fine-tunes. For shipped products on stock open weights, SambaNova is the better fit.
What SambaNova and Thinking Machines do
SambaNova
SambaNova designs its own inference chip, the Reconfigurable Dataflow Unit, and sells fast tokens on large open models through SambaCloud. The RDU maps the model graph onto the chip to cut trips to off-chip memory. A three-tier memory design of SRAM, HBM and bulk DRAM lets one system host very large models and hot swap between several of them in milliseconds. SambaCloud serves models like MiniMax M2.7, DeepSeek, Gemma 4 31B and GPT-OSS 120B, with speeds reported by Artificial Analysis.
Example models: MiniMax M2.7, GPT-OSS 120B
Full SambaNova profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose SambaNova or Thinking Machines?
SambaNova
Choose SambaNova for
- Fast decode on MiniMax M2.7 and GPT-OSS 120B
- Agents that switch between several large models
- Neoclouds adding a premium speed tier
Thinking Machines
Choose Thinking Machines for
- Custom SFT or RL loops on open models
- Fine-tuning large MoE models like Kimi K2.6
- Testing native audio and image input with 1M context
SambaNova vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | MiniMax M2.7, GPT-OSS 120B, DeepSeek | Inkling, Inkling-Small |
| Speed | ~820 tok/s on MiniMax M2.7 (SN50) | Unknown |
| Price | $0.22 in, $0.59 out (GPT-OSS 120B) | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Unknown | LoRA SFT and RL via Tinker |
| Deployment | SambaCloud, racks for neoclouds | Training API, beta serverless (Inkling only) |
| Long context | Up to 192K (MiniMax M2.7) | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between SambaNova and Thinking Machines?
SambaNova serves big open models fast on its own dataflow chip. Thinking Machines is mainly a place to train them, with Tinker for LoRA post-training and a beta API for Inkling.
When should I choose SambaNova over Thinking Machines?
Fast decode on MiniMax M2.7 and GPT-OSS 120B; Agents that switch between several large models; Neoclouds adding a premium speed tier.
When should I choose Thinking Machines over SambaNova?
Custom SFT or RL loops on open models; Fine-tuning large MoE models like Kimi K2.6; Testing native audio and image input with 1M context.
Is SambaNova or Thinking Machines cheaper?
SambaNova: $0.22 in, $0.59 out (GPT-OSS 120B). Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, SambaNova or Thinking Machines?
SambaNova: Up to 192K (MiniMax M2.7). Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs SambaNova
OpenAI vs SambaNova
Anthropic vs SambaNova
Google Vertex AI vs SambaNova
Amazon Bedrock vs SambaNova
Together AI vs SambaNova
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.