We raised $5.1M for long-running agents.
vs

Subconscious vs Thinking Machines

Thinking Machines helps teams train their own open model with Tinker. Subconscious serves open models fast and cheap on long agent traces. One shapes the weights, the other runs them.

By The Subconscious Team · Updated

Subconscious vs Thinking Machines: key differences

These two sit at different points in the open-model lifecycle. Thinking Machines sells Tinker, a post-training API with four low-level calls that lets teams write their own SFT or RL loops on models like Kimi K2.6, GLM-5.3, Qwen3.5 and its own Inkling family, using LoRA adapters. Serving is secondary: a beta serverless API covers only Inkling and Inkling-Small, and the OpenAI-compatible checkpoint endpoint is scoped to testing and low internal traffic. Subconscious is the opposite. It is an inference runtime for long-horizon agents, serving GLM 5.3 and DeepSeek V4.1 Flash on its managed API, with dedicated or on-prem deployments for nearly any open model. Against open models on standard inference it delivers 2x faster task completion, a 5M+ effective context and 50 to 80% lower cost.

Context and billing split them further. Inkling reaches 1M tokens, but Tinker training runs at 32K to 256K depending on the model, and billing meters prefill, sample and train tokens separately. Subconscious prunes the KV cache and bills only tokens processed after compression, which favors agent traces past 200K tokens, and it scores neutral to 10% better on agentic benchmarks. Thinking Machines wins whenever the goal is a custom model: RL on a large MoE base, full control of the training loop, or Apache 2.0 weights with native image and audio input. A team could reasonably post-train on Tinker, then serve the result on a Subconscious dedicated deployment for long agent runs.

What Subconscious and Thinking Machines do

Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

Example models: GLM 5.3, DeepSeek V4.1 Flash

Full Subconscious profile

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

Should you choose Subconscious or Thinking Machines?

Subconscious

Choose Subconscious for

  • Production coding agents running past 200K tokens
  • Long traces billed on processed tokens after compression
  • Dedicated or on-prem serving of an open model

Thinking Machines

Choose Thinking Machines for

  • Custom SFT or RL loops without managing GPU clusters
  • LoRA post-training on large MoE bases like Kimi K2.6
  • Apache 2.0 Inkling models with image and audio input

Subconscious vs Thinking Machines at a glance

AttributeSubconsciousThinking Machines
Model accessOpen weightsOpen weights
Flagship modelsGLM 5.3, DeepSeek V4.1 FlashInkling, Inkling-Small
Speed2x faster task completionUnknown
Price50–80% lower cost; billed on processed tokensPer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out
CustomizationMarathon post-trained variantsLoRA SFT and RL via Tinker
DeploymentManaged API, dedicated, on-premTraining API, beta serverless (Inkling only)
Long context5M+ effective contextInkling up to 1M; Tinker 32K–256K

Frequently asked questions

What is the difference between Subconscious and Thinking Machines?

Thinking Machines helps teams train their own open model with Tinker. Subconscious serves open models fast and cheap on long agent traces. One shapes the weights, the other runs them.

When should I choose Subconscious over Thinking Machines?

Production coding agents running past 200K tokens; Long traces billed on processed tokens after compression; Dedicated or on-prem serving of an open model.

When should I choose Thinking Machines over Subconscious?

Custom SFT or RL loops without managing GPU clusters; LoRA post-training on large MoE bases like Kimi K2.6; Apache 2.0 Inkling models with image and audio input.

Is Subconscious or Thinking Machines cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

Which has more context, Subconscious or Thinking Machines?

Subconscious: 5M+ effective context. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

Related comparisons

Run your longest agent traces on Subconscious

Point the OpenAI or Anthropic SDK, or the coding agent you already use, at Subconscious. Keep Thinking Machines for the work it does best and send the long runs to us.