We raised $5.1M for long-running agents.
vs

Crusoe vs StepFun

StepFun is a Shanghai model lab with Apache 2.0 multimodal weights and a first-party API. Crusoe is a US infrastructure cloud serving other labs' open models.

By The Subconscious Team · Updated

Crusoe vs StepFun: key differences

StepFun builds models; Crusoe hosts them. StepFun's Step 3.7 Flash is a 198B mixture-of-experts vision-language model with 11B active parameters, 256K context, selectable reasoning and tool use, released under Apache 2.0. Its own API prices it at $0.20 in and $1.15 out per million at about 128 tokens per second, and OpenRouter carries it too. Crusoe's serverless catalog lists DeepSeek, GLM, Kimi, Gemma, gpt-oss and Nemotron from $0.05 in and $0.20 out, and does not list StepFun models. Crusoe's pitch is its engine: a cluster-wide KV cache that it claims gives up to 9.9x faster time to first token than vLLM on prefix-heavy work.

Location and support weigh heavily. StepFun's first-party inference is hosted in China with thin Western distribution and support, which many US and EU buyers cannot accept. Its open weights run on vLLM and SGLang, so a team could self-host Step 3.7 Flash on rented GPUs, including Crusoe's H100s at $3.90 per GPU-hour. Crusoe adds managed LoRA fine-tuning, dedicated deployments with SLAs and large clusters. StepFun's advantage is a cheap multimodal model that reads images and video with long context. It does trail frontier models on hard multimodal reasoning benchmarks.

What Crusoe and StepFun do

Crusoe

Crusoe started in 2018 turning wasted natural gas into power for computing and has since become a vertically integrated AI infrastructure company: it sources energy, builds data centers and rents GPUs through Crusoe Cloud. It designed and built the Abilene, Texas campus behind the OpenAI and Oracle Stargate project, planned at 1.2 GW, and in March 2026 announced an adjacent 900 MW campus for Microsoft. On September 17, 2026 it closed the first part of a $3.9B Series F at a $30.9B post-money valuation, and it reports over 6 GW of contracted capacity. Crusoe Cloud lists GB200 NVL72, B200 and AMD MI355X by quote, with H100 at $3.90 and H200 at $4.29 per GPU-hour on demand.

Example models: DeepSeek V4 Pro, GLM 5.3, Kimi K2.6

Full Crusoe profile

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Should you choose Crusoe or StepFun?

Crusoe

Choose Crusoe for

  • Managed serving with US data centers
  • Fine-tuning and dedicated deployments
  • GPU capacity for self-hosting open weights

StepFun

Choose StepFun for

  • Low-cost image and video understanding
  • Apache 2.0 weights with small active size
  • Multimodal agents that need 256K context

Crusoe vs StepFun at a glance

AttributeCrusoeStepFun
Model accessOpen weightsOpen (Apache 2.0) and API models
Flagship modelsDeepSeek V4, GLM 5.3, Kimi K2.6, Nemotron 3Step 3.7 Flash, Step3
SpeedUp to 9.9x faster TTFT vs vLLM (vendor claim)~128 tok/s on Step 3.7 Flash
Price$0.05–$1.74 in, $0.20–$4.40 out per 1M$0.20 in, $1.15 out (Step 3.7 Flash)
CustomizationServerless LoRA fine-tuningOpen weights to fine-tune
DeploymentServerless, self-serve and tailored dedicated, raw GPUsFirst-party API, OpenRouter
Long contextVaries by model; cluster-wide KV cache256K

Frequently asked questions

What is the difference between Crusoe and StepFun?

StepFun is a Shanghai model lab with Apache 2.0 multimodal weights and a first-party API. Crusoe is a US infrastructure cloud serving other labs' open models.

When should I choose Crusoe over StepFun?

Managed serving with US data centers; Fine-tuning and dedicated deployments; GPU capacity for self-hosting open weights.

When should I choose StepFun over Crusoe?

Low-cost image and video understanding; Apache 2.0 weights with small active size; Multimodal agents that need 256K context.

Is Crusoe or StepFun cheaper?

Crusoe: $0.05–$1.74 in, $0.20–$4.40 out per 1M. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

Which has more context, Crusoe or StepFun?

Crusoe: Varies by model; cluster-wide KV cache. StepFun: 256K.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.