Alibaba Cloud vs Thinking Machines
Alibaba Cloud serves a closed Qwen 3.8-Max inside a full public cloud. Thinking Machines trains open Qwen3.5 and other bases through Tinker.
By The Subconscious Team · Updated
Alibaba Cloud vs Thinking Machines: key differences
Alibaba plays both sides of Qwen. Its flagship Qwen 3.8-Max is closed, with text, image and video input, 1M context and built-in web search at $2 in and $6 out internationally, and it lacks fine-tuning and batch support. Smaller Qwen models ship as open weights. Thinking Machines trains on that open side: Tinker lists Qwen3.5 among its bases alongside Kimi K2.6, GLM-5.3 and Inkling, and lets teams write SFT or RL loops with LoRA adapters while the lab runs the GPUs. Its own Inkling is Apache 2.0, 975B parameters with 41B active, with image and audio input and 1M context, at $1.00 in and $4.05 out on a beta serverless API.
Alibaba is the more complete platform. Model Studio adds batch at half price on eligible models, caching, a free 1M token quota per model for 90 days, regional deployment including the EU, and a full hyperscale cloud underneath. Its price sheet is confusing, with region scopes and rotating promotions. Thinking Machines is narrow: serving covers only Inkling, and checkpoint sampling is for testing and low internal traffic. Multilingual and Asia-market products that need a strong hosted flagship fit Alibaba. Teams that want to specialize an open Qwen base, which the Max tier cannot do, fit Tinker.
What Alibaba Cloud and Thinking Machines do
Alibaba Cloud
Alibaba Cloud serves the Qwen model family through Model Studio, its managed AI platform. The flagship Qwen 3.8-Max takes text, image and video input with a 1M token context, function calling, structured outputs and built-in web search. International pricing is $2 in and $6 out per million tokens, with implicit cache hits at $0.25. Deployments in China and some global regions list lower, at $1.65 in and about $4.95 out, and Alibaba often runs limited-time discounts, including night-time cuts of up to 80% on Qwen 3.7-Max.
Example models: Qwen 3.8-Max, Qwen 3.7-Max
Full Alibaba Cloud profileThinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileShould you choose Alibaba Cloud or Thinking Machines?
Alibaba Cloud
Choose Alibaba Cloud for
- Multilingual and Asia-market products on Qwen
- EU regional deployment inside a full cloud
- Video input on a hosted flagship
Thinking Machines
Choose Thinking Machines for
- Fine-tuning open Qwen3.5 where Max allows none
- Custom RL loops without cluster setup
- Apache 2.0 Inkling with audio input
Alibaba Cloud vs Thinking Machines at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Closed Max; open smaller Qwen | Open weights |
| Flagship models | Qwen 3.8-Max, Qwen 3.7-Max | Inkling, Inkling-Small |
| Speed | ~40 tok/s on Qwen 3.8-Max | Unknown |
| Price | $2 in, $6 out international | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | No fine-tuning on Max | LoRA SFT and RL via Tinker |
| Deployment | Model Studio on Alibaba Cloud | Training API, beta serverless (Inkling only) |
| Long context | 1M (Qwen 3.8-Max) | Inkling up to 1M; Tinker 32K–256K |
Frequently asked questions
What is the difference between Alibaba Cloud and Thinking Machines?
Alibaba Cloud serves a closed Qwen 3.8-Max inside a full public cloud. Thinking Machines trains open Qwen3.5 and other bases through Tinker.
When should I choose Alibaba Cloud over Thinking Machines?
Multilingual and Asia-market products on Qwen; EU regional deployment inside a full cloud; Video input on a hosted flagship.
When should I choose Thinking Machines over Alibaba Cloud?
Fine-tuning open Qwen3.5 where Max allows none; Custom RL loops without cluster setup; Apache 2.0 Inkling with audio input.
Is Alibaba Cloud or Thinking Machines cheaper?
Alibaba Cloud: $2 in, $6 out international. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.
Which has more context, Alibaba Cloud or Thinking Machines?
Alibaba Cloud: 1M (Qwen 3.8-Max). Thinking Machines: Inkling up to 1M; Tinker 32K–256K.
Related comparisons
Subconscious vs Alibaba Cloud
OpenAI vs Alibaba Cloud
Anthropic vs Alibaba Cloud
Google Vertex AI vs Alibaba Cloud
Amazon Bedrock vs Alibaba Cloud
Together AI vs Alibaba Cloud
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.