# Thinking Machines vs StepFun

> Two labs shipping open-weight multimodal MoE models under Apache 2.0. StepFun's Step 3.7 Flash is small and cheap. Thinking Machines' Inkling is far larger and comes with a training API.

Canonical: https://www.subconscious.dev/compare/thinking-machines-vs-stepfun · By The Subconscious Team · Updated September 30, 2026

## How they compare

The models make the clearest comparison. StepFun's Step 3.7 Flash is a 198B MoE with 11B active, reading images and video with 256K context, priced at $0.20 in and $1.15 out on StepFun's API. Inkling is a 975B MoE with 41B active and 1M context, taking text, image and audio, at $1.00 in and $4.05 out in beta. Inkling-Small, at 276B total and 12B active, sits closer to Step 3.7 Flash in active size. Both labs release under Apache 2.0, so either can be self-hosted, and Step weights run on vLLM and SGLang.

Distribution and tooling split them. StepFun serves a first-party OpenAI-compatible API hosted in China and is on OpenRouter, but it has thin Western support. Thinking Machines is San Francisco based and offers Tinker, where teams run LoRA SFT or RL on Inkling and other open bases like Qwen3.5 and Kimi K2.6. StepFun offers open weights to fine-tune but no managed training API in its listing. StepFun also builds speech and audio models and trails frontier models on hard multimodal reasoning. Thinking Machines' serving is beta and limited to its two models, so production traffic may need another host either way.

## What each one does

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

### StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

## Which is best, and when

### Choose Thinking Machines for

- Managed RL and SFT on its own open models
- Audio input and 1M context
- A US-based lab for procurement

### Choose StepFun for

- Cheap image and video understanding
- Small-active models for cheap self-hosting
- A mature first-party API with OpenRouter reach

## At a glance

| Attribute | Thinking Machines | StepFun |
|---|---|---|
| Model access | Open weights | Open (Apache 2.0) and API models |
| Flagship models | Inkling, Inkling-Small | Step 3.7 Flash, Step3 |
| Speed | - | ~128 tok/s on Step 3.7 Flash |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | $0.20 in, $1.15 out (Step 3.7 Flash) |
| Customization | LoRA SFT and RL via Tinker | Open weights to fine-tune |
| Deployment | Training API, beta serverless (Inkling only) | First-party API, OpenRouter |
| Long context | Inkling up to 1M; Tinker 32K–256K | 256K |

## FAQ

### What is the difference between Thinking Machines and StepFun?

Two labs shipping open-weight multimodal MoE models under Apache 2.0. StepFun's Step 3.7 Flash is small and cheap. Thinking Machines' Inkling is far larger and comes with a training API.

### When should I choose Thinking Machines over StepFun?

Managed RL and SFT on its own open models; Audio input and 1M context; A US-based lab for procurement.

### When should I choose StepFun over Thinking Machines?

Cheap image and video understanding; Small-active models for cheap self-hosting; A mature first-party API with OpenRouter reach.

### Is Thinking Machines or StepFun cheaper?

Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

### Which has more context, Thinking Machines or StepFun?

Thinking Machines: Inkling up to 1M; Tinker 32K–256K. StepFun: 256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md), [Subconscious vs StepFun](https://www.subconscious.dev/compare/subconscious-vs-stepfun.md).

Full profiles: [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md), [StepFun](https://www.subconscious.dev/providers/stepfun.md).
