# Subconscious vs Thinking Machines

> Thinking Machines helps teams train their own open model with Tinker. Subconscious serves open models fast and cheap on long agent traces. One shapes the weights, the other runs them.

Canonical: https://www.subconscious.dev/compare/subconscious-vs-thinking-machines · By The Subconscious Team · Updated September 30, 2026

## How they compare

These two sit at different points in the open-model lifecycle. Thinking Machines sells Tinker, a post-training API with four low-level calls that lets teams write their own SFT or RL loops on models like Kimi K2.6, GLM-5.3, Qwen3.5 and its own Inkling family, using LoRA adapters. Serving is secondary: a beta serverless API covers only Inkling and Inkling-Small, and the OpenAI-compatible checkpoint endpoint is scoped to testing and low internal traffic. Subconscious is the opposite. It is an inference runtime for long-horizon agents, serving GLM 5.3 and DeepSeek V4.1 Flash on its managed API, with dedicated or on-prem deployments for nearly any open model. Against open models on standard inference it delivers 2x faster task completion, a 5M+ effective context and 50 to 80% lower cost.

Context and billing split them further. Inkling reaches 1M tokens, but Tinker training runs at 32K to 256K depending on the model, and billing meters prefill, sample and train tokens separately. Subconscious prunes the KV cache and bills only tokens processed after compression, which favors agent traces past 200K tokens, and it scores neutral to 10% better on agentic benchmarks. Thinking Machines wins whenever the goal is a custom model: RL on a large MoE base, full control of the training loop, or Apache 2.0 weights with native image and audio input. A team could reasonably post-train on Tinker, then serve the result on a Subconscious dedicated deployment for long agent runs.

## What each one does

### Subconscious

Subconscious is an MIT CSAIL spinout in Kendall Square that builds inference for long-horizon agents, the workloads where a single trace runs past 200K tokens and often into the millions. Its runtime drops in as a replacement for vLLM or SGLang. Instead of rereading an ever-growing context on every step, it prunes the KV cache and preserves suffix state, and Subconscious co-designs the runtime with post-trained model variants it calls Marathon. Against open models on standard inference, Subconscious delivers 2x faster task completion, delivers a 5M+ effective context window, cuts cost 50% and up to 80%, and scores neutral to 10% better on agentic benchmarks.

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

## Which is best, and when

### Choose Subconscious for

- Production coding agents running past 200K tokens
- Long traces billed on processed tokens after compression
- Dedicated or on-prem serving of an open model

### Choose Thinking Machines for

- Custom SFT or RL loops without managing GPU clusters
- LoRA post-training on large MoE bases like Kimi K2.6
- Apache 2.0 Inkling models with image and audio input

## At a glance

| Attribute | Subconscious | Thinking Machines |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | GLM 5.3, DeepSeek V4.1 Flash | Inkling, Inkling-Small |
| Speed | 2x faster task completion | - |
| Price | 50–80% lower cost; billed on processed tokens | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out |
| Customization | Marathon post-trained variants | LoRA SFT and RL via Tinker |
| Deployment | Managed API, dedicated, on-prem | Training API, beta serverless (Inkling only) |
| Long context | 5M+ effective context | Inkling up to 1M; Tinker 32K–256K |

## FAQ

### What is the difference between Subconscious and Thinking Machines?

Thinking Machines helps teams train their own open model with Tinker. Subconscious serves open models fast and cheap on long agent traces. One shapes the weights, the other runs them.

### When should I choose Subconscious over Thinking Machines?

Production coding agents running past 200K tokens; Long traces billed on processed tokens after compression; Dedicated or on-prem serving of an open model.

### When should I choose Thinking Machines over Subconscious?

Custom SFT or RL loops without managing GPU clusters; LoRA post-training on large MoE bases like Kimi K2.6; Apache 2.0 Inkling models with image and audio input.

### Is Subconscious or Thinking Machines cheaper?

Subconscious: 50–80% lower cost; billed on processed tokens. Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. The cheaper choice depends on the model and workload.

### Which has more context, Subconscious or Thinking Machines?

Subconscious: 5M+ effective context. Thinking Machines: Inkling up to 1M; Tinker 32K–256K.

## Try Subconscious

Subconscious speaks the OpenAI and Anthropic API formats. Base URL: https://api.subconscious.dev/v1. Docs: https://docs.subconscious.dev. Get an API key: https://platform.subconscious.dev/signin. Agent guide: https://www.subconscious.dev/agents.md.

Full profiles: [Subconscious](https://www.subconscious.dev/providers/subconscious.md), [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md).
