vs

Modal vs StreamLake

StreamLake is Kuaishou's cloud for the proprietary KAT-Coder models, with data resident in China. Modal is serverless GPU compute for your own models. A closed coding model versus open infrastructure.

By The Subconscious Team · Updated

Modal vs StreamLake: key differences

StreamLake sells model-as-a-service inference and bare metal from Kuaishou, the company behind the Kling video models. Its headline is KAT-Coder-Pro V2.5, a proprietary agentic coding model sold per token or through a KwaiKAT Coding Plan, with a Claude-protocol proxy for Claude Code. Because it is proprietary, it only runs on StreamLake. Modal hosts no models of its own. It gives developers serverless GPUs for Python, billed per second, for open models or custom code. StreamLake also rents bare metal, but mainly to Chinese internet businesses.

The two answer different needs. A developer who wants a cheap subscription coding model inside Claude Code can use StreamLake directly, as long as data residency in China and yuan-first pricing are acceptable. Modal suits teams that want to run an open coding model of their choice, fine-tuned or not, with control over where it runs. Modal's list prices look low, but non-preemptible US production runs about 3.75x list, near $14.81 an hour for an H100.

What Modal and StreamLake do

Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper

Full Modal profile

StreamLake

StreamLake is the AI cloud brand of Kuaishou, the Chinese short-video company behind the Kling video models. It sells model-as-a-service inference and bare-metal compute to internet businesses, drawing on the infrastructure Kuaishou built to serve video at massive scale. Its developer site offers APIs, SDKs and integration guides aimed at taking teams from testing to production.

Example models: KAT-Coder-Pro V2.5, KAT-Coder-Air

Full StreamLake profile

Should you choose Modal or StreamLake?

Modal

Choose Modal for

  • Self-hosting open coding models with full control.
  • Custom GPU jobs billed by the second.
  • Agent sandboxes next to inference.

StreamLake

Choose StreamLake for

  • A subscription plan for agentic coding.
  • KAT-Coder inside Claude Code via a proxy.
  • Domestic MaaS and bare metal for Chinese businesses.

Modal vs StreamLake at a glance

AttributeModalStreamLake
Model accessBring your own weightsProprietary coding models
Flagship modelsNone hostedKAT-Coder-Pro V2.5, KAT-Coder-Air
Speed~1s container bootUnknown
PricePer second; H100 $3.95/hr listPer token or KwaiKAT Coding Plan
CustomizationRun any training codeUnknown
DeploymentServerless GPU containersMaaS API, bare metal
Long contextDepends on the model you deployUnknown

Frequently asked questions

What is the difference between Modal and StreamLake?

StreamLake is Kuaishou's cloud for the proprietary KAT-Coder models, with data resident in China. Modal is serverless GPU compute for your own models. A closed coding model versus open infrastructure.

When should I choose Modal over StreamLake?

Self-hosting open coding models with full control; Custom GPU jobs billed by the second; Agent sandboxes next to inference.

When should I choose StreamLake over Modal?

A subscription plan for agentic coding; KAT-Coder inside Claude Code via a proxy; Domestic MaaS and bare metal for Chinese businesses.

Is Modal or StreamLake cheaper?

Modal: Per second; H100 $3.95/hr list. StreamLake: Per token or KwaiKAT Coding Plan. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.