# Thinking Machines vs RunInfra

> RunInfra automates deployment, benchmarking GPUs and quantizations to hit a latency target. Thinking Machines automates distributed training while leaving the training logic to you.

Canonical: https://www.subconscious.dev/compare/thinking-machines-vs-runinfra · By The Subconscious Team · Updated September 30, 2026

## How they compare

Both hide infrastructure, at different stages. RunInfra's agent takes a plain-English spec, picks a model, benchmarks it across GPUs from L4 to B200, searches AWQ, GPTQ and FP8 variants and ships an OpenAI-compatible endpoint that scales to zero with cold starts under two seconds. Paid plans accept uploads up to 50 GB, and coding plans start at $10 a month for Claude Code, Codex and similar tools. Thinking Machines' Tinker runs the GPU side of LoRA SFT and RL, while teams write the loop with four calls. It trains larger bases than RunInfra hosts, including Kimi K2.6, GLM-5.3 and Inkling.

RunInfra's hosted library is small and centered on mid-size models like Nemotron 3.5 Lightning 30B and Qwen 3.8 27B, well short of frontier quality. Thinking Machines has the bigger model in Inkling, a 975B MoE with 1M context and image and audio input, but serves it only in beta at $1.00 in and $4.05 out. Its checkpoint endpoint is for testing, so a trained adapter needs another host for production, and RunInfra's upload path is one candidate. Both are young. RunInfra launched in 2026 with little independent benchmarking; Thinking Machines has deep funding and a large Nvidia capacity deal.

## What each one does

### Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

### RunInfra

RunInfra pitches open models built for agents, with two ways in. Its hosted Model APIs serve a small curated library, including Nemotron 3.5 Lightning 30B, Qwen 3.8 27B and Ornith 1.5 35B, behind one key that works with both the OpenAI and Anthropic SDKs. Cached context bills at a discount. Coding plans start at $10 a month with limits that reset every five hours and every week, and they plug into Claude Code, Codex, OpenCode, Cline, Aider and dozens of other agent CLIs.

## Which is best, and when

### Choose Thinking Machines for

- Custom RL and SFT on large open models
- Research teams owning the training loop
- Evaluating a 975B open multimodal model

### Choose RunInfra for

- Auto-benchmarked endpoints for a latency target
- Cheap coding plans for agent CLIs
- Voice pipelines without ML ops staff

## At a glance

| Attribute | Thinking Machines | RunInfra |
|---|---|---|
| Model access | Open weights | Open weights |
| Flagship models | Inkling, Inkling-Small | Nemotron 3.5 Lightning 30B, Qwen 3.8 27B |
| Speed | - | Cold starts under 2s |
| Price | Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out | Coding plans from $10 a month |
| Customization | LoRA SFT and RL via Tinker | Uploads up to 50 GB; auto-quantization |
| Deployment | Training API, beta serverless (Inkling only) | Model APIs, agent-built endpoints |
| Long context | Inkling up to 1M; Tinker 32K–256K | Varies by model |

## FAQ

### What is the difference between Thinking Machines and RunInfra?

RunInfra automates deployment, benchmarking GPUs and quantizations to hit a latency target. Thinking Machines automates distributed training while leaving the training logic to you.

### When should I choose Thinking Machines over RunInfra?

Custom RL and SFT on large open models; Research teams owning the training loop; Evaluating a 975B open multimodal model.

### When should I choose RunInfra over Thinking Machines?

Auto-benchmarked endpoints for a latency target; Cheap coding plans for agent CLIs; Voice pipelines without ML ops staff.

### Is Thinking Machines or RunInfra cheaper?

Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. RunInfra: Coding plans from $10 a month. The cheaper choice depends on the model and workload.

### Which has more context, Thinking Machines or RunInfra?

Thinking Machines: Inkling up to 1M; Tinker 32K–256K. RunInfra: Varies by model.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Thinking Machines](https://www.subconscious.dev/compare/subconscious-vs-thinking-machines.md), [Subconscious vs RunInfra](https://www.subconscious.dev/compare/subconscious-vs-runinfra.md).

Full profiles: [Thinking Machines](https://www.subconscious.dev/providers/thinking-machines.md), [RunInfra](https://www.subconscious.dev/providers/runinfra.md).
