We raised $5.1M for long-running agents.
vs

Thinking Machines vs RunInfra

RunInfra automates deployment, benchmarking GPUs and quantizations to hit a latency target. Thinking Machines automates distributed training while leaving the training logic to you.

By The Subconscious Team · Updated

Thinking Machines vs RunInfra: key differences

Both hide infrastructure, at different stages. RunInfra's agent takes a plain-English spec, picks a model, benchmarks it across GPUs from L4 to B200, searches AWQ, GPTQ and FP8 variants and ships an OpenAI-compatible endpoint that scales to zero with cold starts under two seconds. Paid plans accept uploads up to 50 GB, and coding plans start at $10 a month for Claude Code, Codex and similar tools. Thinking Machines' Tinker runs the GPU side of LoRA SFT and RL, while teams write the loop with four calls. It trains larger bases than RunInfra hosts, including Kimi K2.6, GLM-5.3 and Inkling.

RunInfra's hosted library is small and centered on mid-size models like Nemotron 3.5 Lightning 30B and Qwen 3.8 27B, well short of frontier quality. Thinking Machines has the bigger model in Inkling, a 975B MoE with 1M context and image and audio input, but serves it only in beta at $1.00 in and $4.05 out. Its checkpoint endpoint is for testing, so a trained adapter needs another host for production, and RunInfra's upload path is one candidate. Both are young. RunInfra launched in 2026 with little independent benchmarking; Thinking Machines has deep funding and a large Nvidia capacity deal.

What Thinking Machines and RunInfra do

Thinking Machines

Thinking Machines Lab is the San Francisco lab Mira Murati, formerly CTO of OpenAI, founded in February 2025, with OpenAI co-founder John Schulman as chief scientist. It raised about $2 billion at a $12 billion valuation in July 2025 in a round led by Andreessen Horowitz, and in March 2026 signed a multi-year Nvidia deal for one gigawatt of Vera Rubin capacity. Its main developer product is Tinker, an API for post-training open-weight models that launched in October 2025 and is now generally available. Tinker exposes four low-level calls, forward_backward, optim_step, sample and save_state, so teams write their own supervised or reinforcement learning loops while Thinking Machines runs the distributed GPU work. Training uses LoRA adapters rather than full weight updates.

Example models: Inkling, Inkling-Small, Qwen3.5, Kimi K2.6

Full Thinking Machines profile

RunInfra

RunInfra pitches open models built for agents, with two ways in. Its hosted Model APIs serve a small curated library, including Nemotron 3.5 Lightning 30B, Qwen 3.8 27B and Ornith 1.5 35B, behind one key that works with both the OpenAI and Anthropic SDKs. Cached context bills at a discount. Coding plans start at $10 a month with limits that reset every five hours and every week, and they plug into Claude Code, Codex, OpenCode, Cline, Aider and dozens of other agent CLIs.

Example models: Nemotron 3.5 Lightning 30B, Qwen 3.8 27B

Full RunInfra profile

Should you choose Thinking Machines or RunInfra?

Thinking Machines

Choose Thinking Machines for

  • Custom RL and SFT on large open models
  • Research teams owning the training loop
  • Evaluating a 975B open multimodal model

RunInfra

Choose RunInfra for

  • Auto-benchmarked endpoints for a latency target
  • Cheap coding plans for agent CLIs
  • Voice pipelines without ML ops staff

Thinking Machines vs RunInfra at a glance

AttributeThinking MachinesRunInfra
Model accessOpen weightsOpen weights
Flagship modelsInkling, Inkling-SmallNemotron 3.5 Lightning 30B, Qwen 3.8 27B
SpeedUnknownCold starts under 2s
PricePer 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 outCoding plans from $10 a month
CustomizationLoRA SFT and RL via TinkerUploads up to 50 GB; auto-quantization
DeploymentTraining API, beta serverless (Inkling only)Model APIs, agent-built endpoints
Long contextInkling up to 1M; Tinker 32K–256KVaries by model

Frequently asked questions

What is the difference between Thinking Machines and RunInfra?

RunInfra automates deployment, benchmarking GPUs and quantizations to hit a latency target. Thinking Machines automates distributed training while leaving the training logic to you.

When should I choose Thinking Machines over RunInfra?

Custom RL and SFT on large open models; Research teams owning the training loop; Evaluating a 975B open multimodal model.

When should I choose RunInfra over Thinking Machines?

Auto-benchmarked endpoints for a latency target; Cheap coding plans for agent CLIs; Voice pipelines without ML ops staff.

Is Thinking Machines or RunInfra cheaper?

Thinking Machines: Per 1M tokens by prefill, sample, train; Inkling $1.00 in, $4.05 out. RunInfra: Coding plans from $10 a month. The cheaper choice depends on the model and workload.

Which has more context, Thinking Machines or RunInfra?

Thinking Machines: Inkling up to 1M; Tinker 32K–256K. RunInfra: Varies by model.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.