Step3 is our cutting-edge multimodal reasoning model—built on a Mixture-of-Experts architecture with 321B total parameters and 38B active.
It is designed end-to-end to minimize decoding costs while delivering top-tier performance in vision–language reasoning.
Through the co-design of Multi-Matrix Factorization Attention (MFA) and Attention-FFN Disaggregation (AFD),
Step3 maintains exceptional efficiency across both flagship and low-end accelerators.
We introduce how to use our model at inference stage using transformers library. It is recommended to use python=3.10, torch>=2.1.0, and transformers=4.54.0 as the development environment.We currently only support bf16 inference, and multi-patch for image preprocessing is supported by default. This behavior is aligned with vllm and sglang.
step3 huggingface.co is an AI model on huggingface.co that provides step3's model effect (), which can be used instantly with this stepfun-ai step3 model. huggingface.co supports a free trial of the step3 model, and also provides paid use of the step3. Support call step3 model through api, including Node.js, Python, http.
step3 huggingface.co is an online trial and call api platform, which integrates step3's modeling effects, including api services, and provides a free online trial of step3, you can try step3 online for free by clicking the link below.
stepfun-ai step3 online free url in huggingface.co:
step3 is an open source model from GitHub that offers a free installation service, and any user can find step3 on GitHub to install. At the same time, huggingface.co provides the effect of step3 install, users can directly use step3 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.