AI Agent Hub
Back to models
🤖

StepFun: Step 5 Preview

Open Source stepfun Released 2026-09-20
44.0 / 100 600.0B params 1M context Proprietary

About this model

Step 5 Preview is StepFun’s flagship foundation model for real-world agentic work, announced on September 20, 2026. It is a 600-billion-parameter sparse mixture-of-experts (MoE) design with roughly 27 billion active parameters per token across 92 transformer layers, built around an efficiency-first “Pareto frontier” goal: strong intelligence without proportional compute cost. The model supports a one-million-token context window via sparse grouped-query attention with block-wise token merging, configurable reasoning effort (low through high), parallel tool calling, and strict JSON schema output, with an OpenAI-compatible API and open BF16 weights under the StepFun Community License.

Step 5 Preview is aimed at software engineering, long-horizon agents, professional knowledge work, and finance. Official high-effort evaluations highlight strong results on GPQA Diamond (93.5%), terminal and coding agent suites including Terminal-Bench 2.1 (85.0%) and DeepSWE v1.1 (67.7%), web and security agent benchmarks such as BrowseComp (88.7%) and CyberGym (84.7%), and tool-orchestration sets including Toolathlon-Verified (74.1%). With tools, Humanity’s Last Exam rises to 59.4%. Multimodal inputs (text, image, and video) are native; MMMU-Pro is reported at 76.0% at high effort.

The model is positioned for deployment as an autonomous coding and research agent: sustained multi-hour kernel optimization, automated post-training experiments, and workflows that combine search, code execution, and iterative refinement. Local inference is intended for large GPU clusters (about 1.2 TB VRAM in BF16); vLLM and SGLang are the recommended serving stacks for long context and StepFun’s reasoning parser.

Benchmark Scores

HLE
59.4
DeepSWE
67.7
CyberGym
84.7
MMLU-Pro
76.0
MMMU-Pro
76.0
GDPval-AA
1571.0
BrowseComp
88.7
GPQA-Diamond
93.5
AutomationBench
51.0
Agents-Last-Exam
29.5
Terminal-Bench-2.1
85.0
Toolathlon-Verified
74.1

Technical Specs

  • Parameters: 600.0B
  • Architecture: Sparse Mixture-of-Experts (MoE)
  • Context Window: 1,000,000 tokens
  • Input Modalities: text, image, video

Hardware Requirements

  • VRAM: 1200.0 GB
  • Compute: 8× NVIDIA H100 80GB (tensor parallel, BF16)

Pricing

Input Output Currency
1.00 / 1M tokens 2.70 / 1M tokens USD