Long-running agents deserve better inference.
vs

StepFun vs Luminal

StepFun ships efficient multimodal models, many Apache 2.0. Luminal compiles open models like these into faster GPU code.

By The Subconscious Team · Updated

StepFun vs Luminal: key differences

StepFun, a Shanghai lab, releases models such as Step 3.7 Flash and Step3, many under Apache 2.0, and serves them on its own API at $0.20 in and $1.15 out, with 256K context. First-party inference is China-hosted. Luminal makes no models: its compiler turns open weights into native kernels ahead of time for serverless or on-prem serving.

Teams that want StepFun's models outside China can self-host the Apache weights, and Luminal could serve as the engine. Luminal reports 36K tokens per second on GPT-OSS 120B across 8 H100s but has not published multimodal results, so test before relying on it.

What StepFun and Luminal do

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

Example models: GPT-OSS 120B, Llama 3 8B

Full Luminal profile

Should you choose StepFun or Luminal?

StepFun

Choose StepFun for

  • Efficient multimodal open models
  • Apache 2.0 weights
  • Low first-party prices

Luminal

Choose Luminal for

  • Self-hosting open weights outside China
  • On-prem deployments with custom kernel work and SLAs
  • Maximum throughput per GPU on a self-chosen model

StepFun vs Luminal at a glance

AttributeStepFunLuminal
Model accessOpen (Apache 2.0) and API modelsBring your own weights
Flagship modelsStep 3.7 Flash, Step3No public catalog
Speed~128 tok/s on Step 3.7 Flash36K tok/s on GPT-OSS 120B, 8xH100 (vendor)
Price$0.20 in, $1.15 out (Step 3.7 Flash)Pay per use; rates not published
CustomizationOpen weights to fine-tuneCompiles any PyTorch or HF model
DeploymentFirst-party API, OpenRouterServerless (early access), on-prem license
Long context256KUnknown

Frequently asked questions

What is the difference between StepFun and Luminal?

StepFun ships efficient multimodal models, many Apache 2.0. Luminal compiles open models like these into faster GPU code.

When should I choose StepFun over Luminal?

Efficient multimodal open models; Apache 2.0 weights; Low first-party prices.

When should I choose Luminal over StepFun?

Self-hosting open weights outside China; On-prem deployments with custom kernel work and SLAs; Maximum throughput per GPU on a self-chosen model.

Is StepFun or Luminal cheaper?

StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.