# StepFun vs Luminal

> StepFun ships efficient multimodal models, many Apache 2.0. Luminal compiles open models like these into faster GPU code.

Canonical: https://www.subconscious.dev/compare/stepfun-vs-luminal · By The Subconscious Team · Updated September 30, 2026

## How they compare

StepFun, a Shanghai lab, releases models such as Step 3.7 Flash and Step3, many under Apache 2.0, and serves them on its own API at $0.20 in and $1.15 out, with 256K context. First-party inference is China-hosted. Luminal makes no models: its compiler turns open weights into native kernels ahead of time for serverless or on-prem serving.

Teams that want StepFun's models outside China can self-host the Apache weights, and Luminal could serve as the engine. Luminal reports 36K tokens per second on GPT-OSS 120B across 8 H100s but has not published multimodal results, so test before relying on it.

## What each one does

### StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

### Luminal

Luminal builds an inference compiler. Where vLLM and SGLang interpret a model at runtime, Luminal compiles it ahead of time into native kernels for GPUs and ASICs. Models get lowered to a small graph of 15 primitive ops, and the compiler searches over fusion, tiling, memory and scheduling choices instead of relying on hand-written rules, which it says can find optimizations like Flash Attention on its own. The compiler is open source in Rust under Apache 2.0 or MIT, runs on CUDA and Metal with ROCm on the roadmap, and works as a torch.compile backend.

## Which is best, and when

### Choose StepFun for

- Efficient multimodal open models
- Apache 2.0 weights
- Low first-party prices

### Choose Luminal for

- Self-hosting open weights outside China
- On-prem deployments with custom kernel work and SLAs
- Maximum throughput per GPU on a self-chosen model

## At a glance

| Attribute | StepFun | Luminal |
|---|---|---|
| Model access | Open (Apache 2.0) and API models | Bring your own weights |
| Flagship models | Step 3.7 Flash, Step3 | No public catalog |
| Speed | ~128 tok/s on Step 3.7 Flash | 36K tok/s on GPT-OSS 120B, 8xH100 (vendor) |
| Price | $0.20 in, $1.15 out (Step 3.7 Flash) | Pay per use; rates not published |
| Customization | Open weights to fine-tune | Compiles any PyTorch or HF model |
| Deployment | First-party API, OpenRouter | Serverless (early access), on-prem license |
| Long context | 256K | - |

## FAQ

### What is the difference between StepFun and Luminal?

StepFun ships efficient multimodal models, many Apache 2.0. Luminal compiles open models like these into faster GPU code.

### When should I choose StepFun over Luminal?

Efficient multimodal open models; Apache 2.0 weights; Low first-party prices.

### When should I choose Luminal over StepFun?

Self-hosting open weights outside China; On-prem deployments with custom kernel work and SLAs; Maximum throughput per GPU on a self-chosen model.

### Is StepFun or Luminal cheaper?

StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). Luminal: Pay per use; rates not published. The cheaper choice depends on the model and workload.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs StepFun](https://www.subconscious.dev/compare/subconscious-vs-stepfun.md), [Subconscious vs Luminal](https://www.subconscious.dev/compare/subconscious-vs-luminal.md).

Full profiles: [StepFun](https://www.subconscious.dev/providers/stepfun.md), [Luminal](https://www.subconscious.dev/providers/luminal.md).
