Thinking Machines vs StepFun
Two labs shipping open-weight multimodal MoE models under Apache 2.0. StepFun's Step 3.7 Flash is small and cheap. Thinking Machines' Inkling is far larger and comes with a training API.
By The Subconscious Team · Updated
Thinking Machines vs StepFun: key differences
The models make the clearest comparison. StepFun's Step 3.7 Flash is a 198B MoE with 11B active, reading images and video with 256K context, priced at $0.20 in and $1.15 out on StepFun's API. Inkling is a 975B MoE with 41B active and 1M context, taking text, image and audio, at $1.00 in and $4.05 out in beta. Inkling-Small, at 276B total and 12B active, sits closer to Step 3.7 Flash in active size. Both labs release under Apache 2.0, so either can be self-hosted, and Step weights run on vLLM and SGLang.
Distribution and tooling split them. StepFun serves a first-party OpenAI-compatible API hosted in China and is on OpenRouter, but it has thin Western support. Thinking Machines is San Francisco based and offers Tinker, where teams run LoRA SFT or RL on Inkling and other open bases like Qwen3.5 and Kimi K2.6. StepFun offers open weights to fine-tune but no managed training API in its listing. StepFun also builds speech and audio models and trails frontier models on hard multimodal reasoning. Thinking Machines' serving is beta and limited to its two models, so production traffic may need another host either way.
What Thinking Machines and StepFun do
Thinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileStepFun
StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.
Example models: Step 3.7 Flash, Step3
Full StepFun profileShould you choose Thinking Machines or StepFun?
Thinking Machines
Choose Thinking Machines for
- Managed RL and SFT on its own open models
- Audio input and 1M context
- A US-based lab for procurement
StepFun
Choose StepFun for
- Cheap image and video understanding
- Small-active models for cheap self-hosting
- A mature first-party API with OpenRouter reach
Thinking Machines vs StepFun at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Open (Apache 2.0) and API models |
| Flagship models | Inkling, Inkling-Small | Step 3.7 Flash, Step3 |
| Speed | Unknown | ~128 tok/s on Step 3.7 Flash |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | LoRA SFT and RL via Tinker | Open weights to fine-tune |
| Deployment | Training API, beta serverless (Inkling only) | First-party API, OpenRouter |
| Long context | Inkling up to 1M; Tinker 32K–256K | 256K |
Frequently asked questions
What is the difference between Thinking Machines and StepFun?
Two labs shipping open-weight multimodal MoE models under Apache 2.0. StepFun's Step 3.7 Flash is small and cheap. Thinking Machines' Inkling is far larger and comes with a training API.
When should I choose Thinking Machines over StepFun?
Managed RL and SFT on its own open models; Audio input and 1M context; A US-based lab for procurement.
When should I choose StepFun over Thinking Machines?
Cheap image and video understanding; Small-active models for cheap self-hosting; A mature first-party API with OpenRouter reach.
Is Thinking Machines or StepFun cheaper?
Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.
Which has more context, Thinking Machines or StepFun?
Thinking Machines: Inkling up to 1M; Tinker 32K–256K. StepFun: 256K.
Related comparisons
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Subconscious vs StepFun
OpenAI vs StepFun
Anthropic vs StepFun
Google Vertex AI vs StepFun
Amazon Bedrock vs StepFun
Together AI vs StepFun
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.