We raised $5.1M for long-running agents.
vs

fal vs Thinking Machines

fal hosts 1,000+ image, video and audio generation models. Thinking Machines trains language models through Tinker and ships its own multimodal-input Inkling models.

By The Subconscious Team · Updated

fal vs Thinking Machines: key differences

These barely overlap. fal is a generative media platform: FLUX, Kling, Seedream and hundreds of other models behind a queue API with webhooks, billed per image, per video second or GPU time, with no charge for failed outputs on shared endpoints. It offers LoRA training endpoints for media models and serverless H100s from $1.89 an hour. Thinking Machines works on language models. Tinker lets teams write their own SFT or RL loops on open weights like Qwen3.5, Kimi K2.6 and gpt-oss, billed per million tokens across prefill, sample and train meters. Both use LoRA for customization, but on very different kinds of models.

Multimodality points in opposite directions. Inkling and Inkling-Small accept text, image and audio as input with up to 1M tokens of context, but they output text. fal's models output images, video and audio. A product that needs to understand a screenshot or a voice clip and reason over it would look at Inkling, served in beta at $1.00 in and $4.05 out. A product that needs to generate a clip or a product photo belongs on fal. fal's cold starts on less popular endpoints make latency hard to forecast, while Thinking Machines' serving is beta and limited to two models.

What fal and Thinking Machines do

fal

fal is the go-to inference platform for generative media. It hosts 1,000+ image, video and audio models behind one API, including FLUX, Kling, Seedream and other video models, and new releases often land there before competitors have them. Every model page exposes its schema, a playground and example code. Pricing follows the output: per image or megapixel for images, per second or per clip for video, and GPU time for custom work.

Example models: FLUX, Kling

Full fal profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose fal or Thinking Machines?

fal

Choose fal for

  • Adding image or video generation to an app
  • Prototyping across many media models on one bill
  • Long async renders with webhooks

Thinking Machines

Choose Thinking Machines for

  • RL or SFT post-training of open LLMs
  • Reasoning over image and audio input with 1M context
  • Research teams building task-specialized models

fal vs Thinking Machines at a glance

AttributefalThinking Machines
Model accessHosted media modelsOpen weights
Flagship modelsFLUX, Kling, SeedreamInkling, Inkling-Small
SpeedCold starts on less popular endpointsUnknown
PricePer image, per video second, GPU timePer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationLoRA training endpointsLoRA SFT and RL via Tinker
DeploymentHosted API, serverless GPUsTraining API, beta serverless (Inkling only)
Long contextNot applicableInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between fal and Thinking Machines?

fal hosts 1,000+ image, video and audio generation models. Thinking Machines trains language models through Tinker and ships its own multimodal-input Inkling models.

When should I choose fal over Thinking Machines?

Adding image or video generation to an app; Prototyping across many media models on one bill; Long async renders with webhooks.

When should I choose Thinking Machines over fal?

RL or SFT post-training of open LLMs; Reasoning over image and audio input with 1M context; Research teams building task-specialized models.

Is fal or Thinking Machines cheaper?

fal: Per image, per video second, GPU time. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, fal or Thinking Machines?

fal: Not applicable. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.