# fal vs Thinking Machines

> fal hosts 1,000+ image, video and audio generation models. Thinking Machines trains language models through Tinker and ships its own multimodal-input Inkling models.

Canonical: https://www.subconscious.dev/compare/fal-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

These barely overlap. fal is a generative media platform: FLUX, Kling, Seedream and hundreds of other models behind a queue API with webhooks, billed per image, per video second or GPU time, with no charge for failed outputs on shared endpoints. It offers LoRA training endpoints for media models and serverless H100s from $1.89 an hour. Thinking Machines works on language models. Tinker lets teams write their own SFT or RL loops on open weights like Qwen3.5, Kimi K2.6 and gpt-oss, billed per million tokens across prefill, sample and train meters. Both use LoRA for customization, but on very different kinds of models.

Multimodality points in opposite directions. Inkling and Inkling-Small accept text, image and audio as input with up to 1M tokens of context, but they output text. fal's models output images, video and audio. A product that needs to understand a screenshot or a voice clip and reason over it would look at Inkling, served in beta at $1.00 in and $4.05 out. A product that needs to generate a clip or a product photo belongs on fal. fal's cold starts on less popular endpoints make latency hard to forecast, while Thinking Machines' serving is beta and limited to two models.

## What each one does

### fal

fal is the go-to inference platform for generative media. It hosts 1,000+ image, video and audio models behind one API, including FLUX, Kling, Seedream and other video models, and new releases often land there before competitors have them. Every model page exposes its schema, a playground and example code. Pricing follows the output: per image or megapixel for images, per second or per clip for video, and GPU time for custom work.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose fal for

- Adding image or video generation to an app
- Prototyping across many media models on one bill
- Long async renders with webhooks

### Choose Thinking Machines for

- RL or SFT post-training of open LLMs
- Reasoning over image and audio input with 1M context
- Research teams building task-specialized models

## At a glance

| Attribute | fal | Thinking Machines |
|---|---|---|
| Model access | Hosted media models | Open weights |
| Flagship models | FLUX, Kling, Seedream | Inkling, Inkling-Small |
| Speed | Cold starts on less popular endpoints | - |
| Price | Per image, per video second, GPU time | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | LoRA training endpoints | LoRA SFT and RL via Tinker |
| Deployment | Hosted API, serverless GPUs | Training API, beta serverless (Inkling only) |
| Long context | Not applicable | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between fal and Thinking Machines?

fal hosts 1,000+ image, video and audio generation models. Thinking Machines trains language models through Tinker and ships its own multimodal-input Inkling models.

### When should I choose fal over Thinking Machines?

Adding image or video generation to an app; Prototyping across many media models on one bill; Long async renders with webhooks.

### When should I choose Thinking Machines over fal?

RL or SFT post-training of open LLMs; Reasoning over image and audio input with 1M context; Research teams building task-specialized models.

### Is fal or Thinking Machines cheaper?

fal: Per image, per video second, GPU time. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, fal or Thinking Machines?

fal: Not applicable. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs fal](https://www.subconscious.dev/compare/subconscious-vs-fal.md), [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md).

Full profiles: [fal](https://www.subconscious.dev/providers/fal.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
