StepFun (阶跃星辰)

platform.stepfun.com

StepFun builds open-weight multimodal LLMs like Step 3.7 Flash with native vision and tool use, plus a free tier and OpenAI-compatible API aimed at agent workflows.

Overview

StepFun (阶跃星辰) is a Shanghai-based AI lab whose Step series of models has climbed domestic leaderboards and expanded into multimodal and agent workloads. The flagship Step 3.7 Flash uses a sparse MoE architecture (roughly 198B total / 11B active parameters) with native image and video understanding, tool calling, and a 256K context window tuned for agentic coding and planning. A free quota is available to new platform users, and the API is OpenAI-compatible, so it drops into existing agent stacks.

Beyond text, StepFun ships Step1X-Edit for image editing (style transfer, inpainting, text replacement) and StepAudio for speech synthesis and real-time voice. The lab has also pushed toward on-device and terminal AI with its Step AOS agentic OS. In our evaluation, Step 3.7 Flash stands out for high-throughput, multimodal agent loops where a single call can plan, reason, and emit code or structured output — a good fit for builders who want open-ish, cost-effective intelligence with vision built in.

Key Features

  • Native Multimodal — Step 3.7 Flash understands images and video directly, no separate vision model needed in agent loops.
  • Agent-Optimized — MoE architecture delivers fast, low-latency inference for multi-step tool-use and planning.
  • OpenAI-Compatible API — Works with standard SDKs and popular coding agents out of the box.
  • Open Ecosystem — Open-weight models plus Step1X-Edit and training frameworks on GitHub.
  • Long Context — 256K-token window supports long documents, codebases, and conversation history.

Pricing

PlanPriceWhat’s included
Free$0New-user quota, basic chat and API access
API (pay-as-you-go)Usage-basedStep 3.7 Flash and other models, OpenAI-compatible
StepAudio 2.5 Realtime~¥10/M in (¥2 cached), ¥70/M outReal-time voice, WebSocket API
EnterpriseCustomDedicated capacity and SLAs

Comparison

vs. DeepSeek: Both are Chinese open-weight labs with cheap APIs; DeepSeek has a larger global community, while StepFun leads on native multimodal and agent-tuned models. vs. Kimi: Kimi emphasizes long-context chat and document work; StepFun adds built-in vision and tool-use for agent pipelines. vs. Qwen: Qwen offers the broadest open-model range, but StepFun’s Flash models are tuned specifically for real-time agent performance.

Compare alternatives

Side-by-side with the 3 closest alternatives.

ToolCategoryPricingVisit
StepFun (阶跃星辰) (this) chat, agentsFree $0 Site ↗
DeepSeekchat, codeFrom $0/mo Site ↗
KimichatFree $0/mo · From $0.55/mo Site ↗
Qwen (通义千问)chat, codeFree $0 Site ↗

StepFun builds open-weight multimodal LLMs like Step 3.7 Flash with native vision and tool use, plus a free tier and OpenAI-compatible API aimed at agent workflows.

chatagents
Free $0

Open, low-cost reasoning models with strong math and coding performance. Frontier-level reasoning at low cost Read our hands-on review and compare the top

chatcode
From $0/mo

Moonshot AI's assistant powered by the K2.6 MoE model, built for long-horizon coding, agent swarms, and native multimodal chat with a 262K context window.

chat
Free $0/mo · From $0.55/mo

Alibaba's open-weight Qwen LLM family and Qwen Chat assistant for multilingual conversation, coding, and document Q&A. Genuinely open-weight smaller

chatcode
Free $0
Editor’s Review
4.2/5
Pros
  • +Native multimodal models (text + image + video understanding)
  • +Step 3.7 Flash is built for agents with fast MoE inference
  • +Free starter quota for new developers, OpenAI-compatible API
  • +Open-source models and ecosystem (Step1X-Edit, training frameworks)
Cons
  • Primarily China-based; documentation skews Chinese-first
  • Free tier suits prototyping more than heavy production use
  • Ecosystem and tooling less mature than Western incumbents

StepFun is an increasingly serious open-model lab whose Step 3.7 Flash shines in agent and multimodal tasks at a friendly price. It is a strong pick for builders needing vision-capable, tool-calling models without the cost of frontier APIs, though English-language support and ecosystem maturity still lag.

See all reviews →