vs

TypeSafe AI vs StepFun

StepFun builds cheap multimodal generators; TypeSafe builds a decision model that returns typed answers in about 100ms. Jev can route work to Step 3.7 Flash, not replace it.

By The Subconscious Team · Updated

TypeSafe AI vs StepFun: key differences

StepFun and TypeSafe build different kinds of models. Step 3.7 Flash is a general vision-language model: 198B parameters with 11B active, 256K context, tool use, structured outputs and selectable reasoning levels, at $0.20 in and $1.15 out, with Apache 2.0 weights. TypeSafe's Jev generates nothing. It takes text, evaluates a fixed set of options in one pass and returns a typed answer with calibrated probabilities, usually in about 100ms. StepFun's structured outputs shape generated text into a schema. With Jev, the answer space is the schema, so it cannot return a value outside it.

Input is a clear divider. Jev is text-only for now, while Step 3.7 Flash reads images and video, and StepFun also builds speech models. For a decision about a screenshot or a clip, StepFun is the only one of the two that can see it. For high-volume text decisions like intent classification or ticket routing, TypeSafe claims 40 to 200x the speed of an LLM at a fraction of the cost, with confidence scores that tell software when to escalate. Both have access caveats: Jev is early access, and StepFun's first-party inference is China-hosted with thin Western support.

What TypeSafe AI and StepFun do

TypeSafe AI

TypeSafe AI builds decision models instead of text generators. Founder Diogo Almeida co-invented RLHF and InstructGPT at OpenAI and later worked at Google Brain. After two years in stealth the company released its first System One Model, Jev, in early access. The name nods to Kahneman's fast System 1 thinking, and the model is built for machines to call, not people to chat with.

Example models: Jev, jev-1.13

Full TypeSafe AI profile

StepFun

StepFun is a Shanghai AI lab known for efficient multimodal models, with a mix of proprietary API models and open-weight releases. Its current workhorse, Step 3.7 Flash, came out in May 2026 as a 198B mixture-of-experts vision-language model with only 11B active parameters. It has 256K context, selectable reasoning levels, tool use and structured outputs, and it ships under Apache 2.0. StepFun's own API prices it at $0.20 in and $1.15 out per million tokens, and OpenRouter carries it too.

Example models: Step 3.7 Flash, Step3

Full StepFun profile

Should you choose TypeSafe AI or StepFun?

TypeSafe AI

Choose TypeSafe AI for

  • Text routing and intent classification in about 100ms
  • Decisions that must stay inside a fixed set of labels
  • Escalation logic driven by calibrated confidence

StepFun

Choose StepFun for

  • Decisions and tasks that depend on images or video
  • Open-ended generation with tool use and structured outputs
  • Self-hosting under Apache 2.0

TypeSafe AI vs StepFun at a glance

AttributeTypeSafe AIStepFun
Model accessDecision modelsOpen (Apache 2.0) and API models
Flagship modelsJev, jev-1.13Step 3.7 Flash, Step3
Speed~100ms per call~128 tok/s on Step 3.7 Flash
PriceA fraction of an LLM call$0.20 in, $1.15 out (Step 3.7 Flash)
CustomizationUnknownOpen weights to fine-tune
DeploymentEarly-access APIFirst-party API, OpenRouter
Long contextUnknown256K

Frequently asked questions

What is the difference between TypeSafe AI and StepFun?

StepFun builds cheap multimodal generators; TypeSafe builds a decision model that returns typed answers in about 100ms. Jev can route work to Step 3.7 Flash, not replace it.

When should I choose TypeSafe AI over StepFun?

Text routing and intent classification in about 100ms; Decisions that must stay inside a fixed set of labels; Escalation logic driven by calibrated confidence.

When should I choose StepFun over TypeSafe AI?

Decisions and tasks that depend on images or video; Open-ended generation with tool use and structured outputs; Self-hosting under Apache 2.0.

Is TypeSafe AI or StepFun cheaper?

TypeSafe AI: A fraction of an LLM call. StepFun: $0.20 in, $1.15 out (Step 3.7 Flash). The cheaper choice depends on the model and workload.

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.