# RunInfra

> Open models for agents, plus an agent that benchmarks and builds deployments for you.

Canonical: https://www.subconscious.dev/providers/runinfra · By The Subconscious Team · Updated September 30, 2026

- Founded: 2026
- Example models: Nemotron 3.5 Lightning 30B, Qwen 3.8 27B
- Website: https://runinfra.ai

## Overview

RunInfra pitches open models built for agents, with two ways in. Its hosted Model APIs serve a small curated library, including Nemotron 3.5 Lightning 30B, Qwen 3.8 27B and Ornith 1.5 35B, behind one key that works with both the OpenAI and Anthropic SDKs. Cached context bills at a discount. Coding plans start at $10 a month with limits that reset every five hours and every week, and they plug into Claude Code, Codex, OpenCode, Cline, Aider and dozens of other agent CLIs.

The second product is an agent that builds deployments for you. You describe the endpoint in plain English, like a low-latency chatbot on Llama 3.1 8B, and RunInfra picks a model, benchmarks it across GPUs from L4 to B200, searches quantized variants like AWQ, GPTQ and FP8, and applies its own Forge kernels. It then ships an OpenAI-compatible endpoint that scales to zero with cold starts under two seconds, or runs always-on for production. Paid plans accept custom uploads up to 50 GB in SafeTensors, GGUF or ONNX, and pipelines can chain models such as Whisper into an LLM into a TTS voice.

## Upsides

- Cheap flat-rate coding plans that work with most popular agent harnesses.
- Automated benchmarking and quantization that picks the cheapest config meeting a latency target.

## Core use cases

- Developers who want a low-cost open model inside Claude Code or Codex-style tools.
- Small teams deploying a tuned open model or voice pipeline without ML ops staff.

## Downsides

- Hosted library is tiny and centered on mid-size models, far from frontier quality.
- Young company with little independent benchmarking or enterprise track record.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open weights |
| Flagship models | Nemotron 3.5 Lightning 30B, Qwen 3.8 27B |
| Speed | Cold starts under 2s |
| Price | Coding plans from $10 a month |
| Customization | Uploads up to 50 GB; auto-quantization |
| Deployment | Model APIs, agent-built endpoints |
| Long context | Varies by model |

## FAQ

### What is RunInfra?

RunInfra pitches open models built for agents, with two ways in. Its hosted Model APIs serve a small curated library, including Nemotron 3.5 Lightning 30B, Qwen 3.8 27B and Ornith 1.5 35B, behind one key that works with both the OpenAI and Anthropic SDKs. Cached context bills at a discount. Coding plans start at $10 a month with limits that reset every five hours and every week, and they plug into Claude Code, Codex, OpenCode, Cline, Aider and dozens of other agent CLIs.

### What is RunInfra best for?

Developers who want a low-cost open model inside Claude Code or Codex-style tools; Small teams deploying a tuned open model or voice pipeline without ML ops staff.

### How much does RunInfra cost?

RunInfra pricing at a glance: Coding plans from $10 a month. Rates change often, so check RunInfra's pricing page before committing.

### How much context does RunInfra support?

RunInfra's long-context support: Varies by model.

### What are the downsides of RunInfra?

Hosted library is tiny and centered on mid-size models, far from frontier quality; Young company with little independent benchmarking or enterprise track record.

### What are the best alternatives to RunInfra?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with RunInfra on this site.

## Comparisons

- [Subconscious vs RunInfra](https://www.subconscious.dev/compare/subconscious-vs-runinfra.md)
- [OpenAI vs RunInfra](https://www.subconscious.dev/compare/openai-vs-runinfra.md)
- [Anthropic vs RunInfra](https://www.subconscious.dev/compare/anthropic-vs-runinfra.md)
- [Google Vertex AI vs RunInfra](https://www.subconscious.dev/compare/google-vertex-vs-runinfra.md)
- [Amazon Bedrock vs RunInfra](https://www.subconscious.dev/compare/aws-bedrock-vs-runinfra.md)
- [Together AI vs RunInfra](https://www.subconscious.dev/compare/together-ai-vs-runinfra.md)
- [Fireworks AI vs RunInfra](https://www.subconscious.dev/compare/fireworks-vs-runinfra.md)
- [Baseten vs RunInfra](https://www.subconscious.dev/compare/baseten-vs-runinfra.md)
- [Groq vs RunInfra](https://www.subconscious.dev/compare/groq-vs-runinfra.md)
- [Cerebras vs RunInfra](https://www.subconscious.dev/compare/cerebras-vs-runinfra.md)
- [DeepInfra vs RunInfra](https://www.subconscious.dev/compare/deepinfra-vs-runinfra.md)
- [Modal vs RunInfra](https://www.subconscious.dev/compare/modal-vs-runinfra.md)
- [xAI vs RunInfra](https://www.subconscious.dev/compare/xai-vs-runinfra.md)
- [DeepSeek vs RunInfra](https://www.subconscious.dev/compare/deepseek-vs-runinfra.md)
- [Moonshot AI vs RunInfra](https://www.subconscious.dev/compare/moonshot-ai-vs-runinfra.md)
- [Z.ai vs RunInfra](https://www.subconscious.dev/compare/z-ai-vs-runinfra.md)
- [Alibaba Cloud vs RunInfra](https://www.subconscious.dev/compare/alibaba-cloud-vs-runinfra.md)
- [Meta vs RunInfra](https://www.subconscious.dev/compare/meta-vs-runinfra.md)
- [SambaNova vs RunInfra](https://www.subconscious.dev/compare/sambanova-vs-runinfra.md)
- [Nebius vs RunInfra](https://www.subconscious.dev/compare/nebius-vs-runinfra.md)
- [fal vs RunInfra](https://www.subconscious.dev/compare/fal-vs-runinfra.md)
- [Novita AI vs RunInfra](https://www.subconscious.dev/compare/novita-ai-vs-runinfra.md)
- [Parasail vs RunInfra](https://www.subconscious.dev/compare/parasail-vs-runinfra.md)
- [Inference.net vs RunInfra](https://www.subconscious.dev/compare/inference-net-vs-runinfra.md)
- [GMI Cloud vs RunInfra](https://www.subconscious.dev/compare/gmi-cloud-vs-runinfra.md)
- [Sail Research vs RunInfra](https://www.subconscious.dev/compare/sail-research-vs-runinfra.md)
- [Morph vs RunInfra](https://www.subconscious.dev/compare/morph-vs-runinfra.md)
- [Relace vs RunInfra](https://www.subconscious.dev/compare/relace-vs-runinfra.md)
- [TypeSafe AI vs RunInfra](https://www.subconscious.dev/compare/typesafe-ai-vs-runinfra.md)
- [StepFun vs RunInfra](https://www.subconscious.dev/compare/stepfun-vs-runinfra.md)
- [Runware vs RunInfra](https://www.subconscious.dev/compare/runware-vs-runinfra.md)
- [StreamLake vs RunInfra](https://www.subconscious.dev/compare/streamlake-vs-runinfra.md)
- [Wafer vs RunInfra](https://www.subconscious.dev/compare/wafer-vs-runinfra.md)
- [RunInfra vs Particle.AI](https://www.subconscious.dev/compare/runinfra-vs-particle-ai.md)

## Sources

- [RunInfra](https://runinfra.ai/)
- [RunInfra pricing](https://runinfra.ai/pricing)
- [RunInfra docs](https://runinfra.mintlify.app/)

Pricing and model lineups change often; figures are a snapshot.
