AI Agent Hub
Back to models
🤖

StepFun: Step 3.7 Flash

Closed Source stepfun
5.6 / 10 262.1K context Proprietary

About this model

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

Benchmark Scores

SWE-Bench-Pro
56.3

Technical Specs

  • Architecture: text+image+video->text
  • Context Window: 262,144 tokens
  • Input Modalities: text

Hardware Requirements

  • API-only (no local hardware needed)

Pricing

Input Output Currency
0.2 / 1M tokens 1.15 / 1M tokens USD