Modal vs Crusoe
Modal is serverless GPU compute for your own Python code, billed by the second. Crusoe offers hosted open models per token plus dedicated and raw GPUs.
By The Subconscious Team · Updated
Modal vs Crusoe: key differences
Modal and Crusoe both rent GPUs, but they meet developers at different layers. Modal has no model catalog or per-token price. A developer decorates a Python function with the hardware it needs, and Modal builds, schedules and autoscales the container and scales it to zero, billing per second from $0.59 an hour for a T4 to $3.95 for an H100 at list. Crusoe sells per-token serverless inference on DeepSeek, GLM, Kimi, Gemma, gpt-oss and Nemotron from $0.05 in and $0.20 out per million, self-serve deployments at $5.50 per H100-hour, and raw H100s at $3.90 on demand. Context on Modal depends on the model you deploy, and on Crusoe it varies by model.
Modal wins on flexibility and bursty load. It runs any training or serving code, covers embeddings, OCR, batch jobs and agent sandboxes, and its free Starter plan renews $30 of credits monthly. The catch is cost at steady load: non-preemptible US production runs about 3.75x list, near $14.81 an hour for an H100, and keeping containers warm turns serverless into always-on. Crusoe wins when a team wants a hosted open model without writing serving code, or steady GPU capacity. Its MemoryAlloy cache, which Crusoe claims gives up to 9.9x faster time to first token versus vLLM on prefix-heavy work, plus LoRA fine-tuning and GB200 or B200 clusters by quote, suit sustained production traffic.
What Modal and Crusoe do
Modal
Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.
Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper
Full Modal profileCrusoe
Crusoe started in 2018 turning wasted natural gas into power for computing and has since become a vertically integrated AI infrastructure company: it sources energy, builds data centers and rents GPUs through Crusoe Cloud. It designed and built the Abilene, Texas campus behind the OpenAI and Oracle Stargate project, planned at 1.2 GW, and in March 2026 announced an adjacent 900 MW campus for Microsoft. On September 17, 2026 it closed the first part of a $3.9B Series F at a $30.9B post-money valuation, and it reports over 6 GW of contracted capacity. Crusoe Cloud lists GB200 NVL72, B200 and AMD MI355X by quote, with H100 at $3.90 and H200 at $4.29 per GPU-hour on demand.
Example models: DeepSeek V4 Pro, GLM 5.3, Kimi K2.6
Full Crusoe profileShould you choose Modal or Crusoe?
Modal vs Crusoe at a glance
| Attribute | ||
|---|---|---|
| Model access | Bring your own weights | Open weights |
| Flagship models | None hosted | DeepSeek V4, GLM 5.3, Kimi K2.6, Nemotron 3 |
| Speed | ~1s container boot | Up to 9.9x faster TTFT vs vLLM (vendor claim) |
| Price | Per second; H100 $3.95/hr list | $0.05–$1.74 in, $0.20–$4.40 out per 1M |
| Customization | Run any training code | Serverless LoRA fine-tuning |
| Deployment | Serverless GPU containers | Serverless, self-serve and tailored dedicated, raw GPUs |
| Long context | Depends on the model you deploy | Varies by model; cluster-wide KV cache |
Frequently asked questions
What is the difference between Modal and Crusoe?
Modal is serverless GPU compute for your own Python code, billed by the second. Crusoe offers hosted open models per token plus dedicated and raw GPUs.
When should I choose Modal over Crusoe?
Spiky GPU jobs like embeddings and transcription; Custom Python serving code that scales to zero; Free monthly credits for experiments.
When should I choose Crusoe over Modal?
Hosted open models with no serving code; Steady production load on dedicated endpoints; Large B200 or GB200 clusters by quote.
Is Modal or Crusoe cheaper?
Modal: Per second; H100 $3.95/hr list. Crusoe: $0.05–$1.74 in, $0.20–$4.40 out per 1M. The cheaper choice depends on the model and workload.
Which has more context, Modal or Crusoe?
Modal: Depends on the model you deploy. Crusoe: Varies by model; cluster-wide KV cache.
Related comparisons
Running long-horizon agents?
If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.