# SambaNova

> Fast decode on large open models, served on its own dataflow chip.

Canonical: https://www.subconscious.dev/providers/sambanova · By The Subconscious Team · Updated September 30, 2026

- Founded: 2017
- Example models: MiniMax M2.7, GPT-OSS 120B
- Website: https://sambanova.ai

## Overview

SambaNova designs its own inference chip, the Reconfigurable Dataflow Unit, and sells fast tokens on large open models through SambaCloud. The RDU maps the model graph onto the chip to cut trips to off-chip memory. A three-tier memory design of SRAM, HBM and bulk DRAM lets one system host very large models and hot swap between several of them in milliseconds. SambaCloud serves models like MiniMax M2.7, DeepSeek, Gemma 4 31B and GPT-OSS 120B, with speeds reported by Artificial Analysis.

The fifth-generation SN50 chip, announced in February 2026, starts shipping in the second half of the year. SambaNova claims 5x the peak speed of an NVIDIA B200 and support for models up to 10 trillion parameters with 10M token contexts, all in a 20 kW air-cooled rack. Its August 2026 pitch leans into what it calls premium inference: GPUs handle prefill, RDUs handle decode, and a SambaRack SN50 runs MiniMax M2.7 near 820 tokens per second in its fastest configuration. SambaNova also sells racks to neoclouds that want to offer a fast tier without replacing their GPU fleet.

## Upsides

- Fast decode on large frontier-scale open models, where Groq and Cerebras have thinner catalogs.
- Millisecond model hot swapping and input caching suit agents that bounce between models.
- Air-cooled racks fit existing data centers.

## Core use cases

- Interactive coding agents and copilots on big open models.
- Neoclouds adding a premium speed tier through disaggregated prefill and decode.

## Downsides

- Smaller public catalog than GPU clouds, and many headline numbers are vendor benchmarks on hardware still ramping.
- Much of the value arrives through hardware sales and partnerships rather than a big self-serve developer platform.

## At a glance

| Attribute | Value |
|---|---|
| Model access | Open weights |
| Flagship models | MiniMax M2.7, GPT-OSS 120B, DeepSeek |
| Speed | ~820 tok/s on MiniMax M2.7 (SN50) |
| Price | $0.22 in, $0.59 out (GPT-OSS 120B) |
| Customization | - |
| Deployment | SambaCloud, racks for neoclouds |
| Long context | Up to 192K (MiniMax M2.7) |

## FAQ

### What is SambaNova?

SambaNova designs its own inference chip, the Reconfigurable Dataflow Unit, and sells fast tokens on large open models through SambaCloud. The RDU maps the model graph onto the chip to cut trips to off-chip memory. A three-tier memory design of SRAM, HBM and bulk DRAM lets one system host very large models and hot swap between several of them in milliseconds. SambaCloud serves models like MiniMax M2.7, DeepSeek, Gemma 4 31B and GPT-OSS 120B, with speeds reported by Artificial Analysis.

### What is SambaNova best for?

Interactive coding agents and copilots on big open models; Neoclouds adding a premium speed tier through disaggregated prefill and decode.

### How much does SambaNova cost?

SambaNova pricing at a glance: $0.22 in, $0.59 out (GPT-OSS 120B). Rates change often, so check SambaNova's pricing page before committing.

### How much context does SambaNova support?

SambaNova's long-context support: Up to 192K (MiniMax M2.7).

### What are the downsides of SambaNova?

Smaller public catalog than GPU clouds, and many headline numbers are vendor benchmarks on hardware still ramping; Much of the value arrives through hardware sales and partnerships rather than a big self-serve developer platform.

### What are the best alternatives to SambaNova?

Common alternatives include Subconscious, OpenAI, Anthropic, Google Vertex AI, Amazon Bedrock. Each has a head-to-head comparison with SambaNova on this site.

## Comparisons

- [Subconscious vs SambaNova](https://www.subconscious.dev/compare/subconscious-vs-sambanova.md)
- [OpenAI vs SambaNova](https://www.subconscious.dev/compare/openai-vs-sambanova.md)
- [Anthropic vs SambaNova](https://www.subconscious.dev/compare/anthropic-vs-sambanova.md)
- [Google Vertex AI vs SambaNova](https://www.subconscious.dev/compare/google-vertex-vs-sambanova.md)
- [Amazon Bedrock vs SambaNova](https://www.subconscious.dev/compare/aws-bedrock-vs-sambanova.md)
- [Together AI vs SambaNova](https://www.subconscious.dev/compare/together-ai-vs-sambanova.md)
- [Fireworks AI vs SambaNova](https://www.subconscious.dev/compare/fireworks-vs-sambanova.md)
- [Baseten vs SambaNova](https://www.subconscious.dev/compare/baseten-vs-sambanova.md)
- [Groq vs SambaNova](https://www.subconscious.dev/compare/groq-vs-sambanova.md)
- [Cerebras vs SambaNova](https://www.subconscious.dev/compare/cerebras-vs-sambanova.md)
- [DeepInfra vs SambaNova](https://www.subconscious.dev/compare/deepinfra-vs-sambanova.md)
- [Modal vs SambaNova](https://www.subconscious.dev/compare/modal-vs-sambanova.md)
- [xAI vs SambaNova](https://www.subconscious.dev/compare/xai-vs-sambanova.md)
- [DeepSeek vs SambaNova](https://www.subconscious.dev/compare/deepseek-vs-sambanova.md)
- [Moonshot AI vs SambaNova](https://www.subconscious.dev/compare/moonshot-ai-vs-sambanova.md)
- [Z.ai vs SambaNova](https://www.subconscious.dev/compare/z-ai-vs-sambanova.md)
- [Alibaba Cloud vs SambaNova](https://www.subconscious.dev/compare/alibaba-cloud-vs-sambanova.md)
- [Meta vs SambaNova](https://www.subconscious.dev/compare/meta-vs-sambanova.md)
- [SambaNova vs Nebius](https://www.subconscious.dev/compare/sambanova-vs-nebius.md)
- [SambaNova vs fal](https://www.subconscious.dev/compare/sambanova-vs-fal.md)
- [SambaNova vs Novita AI](https://www.subconscious.dev/compare/sambanova-vs-novita-ai.md)
- [SambaNova vs Parasail](https://www.subconscious.dev/compare/sambanova-vs-parasail.md)
- [SambaNova vs Inference.net](https://www.subconscious.dev/compare/sambanova-vs-inference-net.md)
- [SambaNova vs GMI Cloud](https://www.subconscious.dev/compare/sambanova-vs-gmi-cloud.md)
- [SambaNova vs Sail Research](https://www.subconscious.dev/compare/sambanova-vs-sail-research.md)
- [SambaNova vs Morph](https://www.subconscious.dev/compare/sambanova-vs-morph.md)
- [SambaNova vs Relace](https://www.subconscious.dev/compare/sambanova-vs-relace.md)
- [SambaNova vs TypeSafe AI](https://www.subconscious.dev/compare/sambanova-vs-typesafe-ai.md)
- [SambaNova vs StepFun](https://www.subconscious.dev/compare/sambanova-vs-stepfun.md)
- [SambaNova vs Runware](https://www.subconscious.dev/compare/sambanova-vs-runware.md)
- [SambaNova vs StreamLake](https://www.subconscious.dev/compare/sambanova-vs-streamlake.md)
- [SambaNova vs Wafer](https://www.subconscious.dev/compare/sambanova-vs-wafer.md)
- [SambaNova vs RunInfra](https://www.subconscious.dev/compare/sambanova-vs-runinfra.md)
- [SambaNova vs Particle.AI](https://www.subconscious.dev/compare/sambanova-vs-particle-ai.md)

## Sources

- [SambaCloud models and context lengths](https://sambanova-systems.mintlify.dev/docs/en/models/sambacloud-models)
- [SN50 announcement](https://sambanova.ai/blog/introducing-the-sn50-rdu-purpose-built-for-agentic-inference)
- [Premium inference, SambaNova](https://sambanova.ai/blog/the-premium-inference-moment-is-here)
- [SambaCloud](https://sambanova.ai/products/sambacloud)

Pricing and model lineups change often; figures are a snapshot.
