Thinking Machines vs Runware
Runware sells low-cost media generation across 300+ priced models. Thinking Machines trains and serves language models, with Inkling reading images and audio but writing only text.
By The Subconscious Team · Updated
Thinking Machines vs Runware: key differences
These are different product categories. Runware's single task-based API covers image, video, audio, 3D and text generation, with images from fractions of a cent and video billed per second, such as Seedance 2.5 at about $0.10 a second at 480p. Its Sonic Inference Engine and Model Lake keep 400K+ models resident, and it rents H100s at $2.76 an hour. It supports fine-tuned diffusion checkpoints. Thinking Machines focuses on language models: Tinker for custom LoRA SFT and RL on open weights, and a beta serverless API for Inkling and Inkling-Small at $1.00 in and $4.05 out for Inkling.
The choice follows the output. For an app that generates images or short video at volume, Runware's pricing and unified schema fit, though outputs expire after seven days by default, so apps need their own storage. Runware calls LLM hosting a side line. For a product that needs to understand an image or audio clip, or train a model on its own data with RL, Thinking Machines is the relevant option. Inkling brings 1M context and Apache 2.0 weights. Neither is a general LLM host for production traffic, since Thinking Machines scopes its checkpoint endpoint to testing.
What Thinking Machines and Runware do
Thinking Machines
Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.
Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6
Full Thinking Machines profileRunware
Runware sells what it calls the lowest-cost API for media generation, and it claims more than 1M developers. One endpoint covers image, video, audio, 3D and text. Every request is a task with the same shape, so switching from a Kling video to a Seedream image mostly means changing the model ID. The published rate sheet lists 300+ priced models, with images from fractions of a cent to a few cents each and video billed per second, like Seedance 2.5 at about $0.10 a second at 480p.
Example models: Seedance 2.5, Qwen-Image-3.0
Full Runware profileShould you choose Thinking Machines or Runware?
Thinking Machines
Choose Thinking Machines for
- RL post-training of open language models
- Image and audio understanding with 1M context
- Research teams building specialized LLMs
Runware
Choose Runware for
- High-volume image and video generation
- One request schema across media types
- Serving community diffusion checkpoints
Thinking Machines vs Runware at a glance
| Attribute | Thinking Machines | |
|---|---|---|
| Model access | Open weights | Hosted media models |
| Flagship models | Inkling, Inkling-Small | Seedance 2.5, Qwen-Image-3.0 |
| Speed | Unknown | Unknown |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | Images from fractions of a cent |
| Customization | LoRA SFT and RL via Tinker | Fine-tuned diffusion checkpoints |
| Deployment | Training API, beta serverless (Inkling only) | Unified API, raw GPUs |
| Long context | Inkling up to 1M; Tinker 32K–256K | Not applicable |
Frequently asked questions
What is the difference between Thinking Machines and Runware?
Runware sells low-cost media generation across 300+ priced models. Thinking Machines trains and serves language models, with Inkling reading images and audio but writing only text.
When should I choose Thinking Machines over Runware?
RL post-training of open language models; Image and audio understanding with 1M context; Research teams building specialized LLMs.
When should I choose Runware over Thinking Machines?
High-volume image and video generation; One request schema across media types; Serving community diffusion checkpoints.
Is Thinking Machines or Runware cheaper?
Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. Runware: Images from fractions of a cent. The cheaper choice depends on the model and workload.
Which has more context, Thinking Machines or Runware?
Thinking Machines: Inkling up to 1M; Tinker 32K–256K. Runware: Not applicable.
Related comparisons
Subconscious vs Thinking Machines
OpenAI vs Thinking Machines
Anthropic vs Thinking Machines
Google Vertex AI vs Thinking Machines
Amazon Bedrock vs Thinking Machines
Together AI vs Thinking Machines
Subconscious vs Runware
OpenAI vs Runware
Anthropic vs Runware
Google Vertex AI vs Runware
Amazon Bedrock vs Runware
Together AI vs Runware
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.