This is the more aggressively pruned Step 3.7 Flash NVFP4 checkpoint. It removes about 50.21B parameters by keeping 212 of 288 routed experts per MoE layer.
The goal is a smaller Step 3.7 Flash derivative for fit and serving experiments. It should be treated as an experimental compression artifact until load, generation, coherence, and benchmark evidence are complete.
Intended serving path: vLLM/SGLang/Transformers paths that support Step 3.7 Flash remote code and ModelOpt NVFP4
Status: private experimental checkpoint; validate fit, generation, and benchmark behavior before production use
How the REAP checkpoint was made
REAP is a one-shot MoE compression method that uses router-weighted expert activation observations to rank experts by practical usefulness. The observation pass records per-layer routed expert activity under calibration prompts, then each MoE layer is pruned independently.
For this checkpoint:
Start from the Step 3.7 Flash NVFP4 checkpoint.
Run calibration data through the model and record router/expert activation observations.
Aggregate expert scores with the
reap_score
metric.
Keep the top routed experts per MoE layer.
Rewrite the checkpoint with pruned expert tensors and updated routing metadata.
Embeddings, attention blocks, normalization, router gates, shared experts, selected routed experts, vision components, tokenizer files, and generation/config files are preserved. The
prune_summary.json
and
layer_expert_metrics.parquet
files in this repo contain the exact pruning map and expert metrics.
Calibration evidence
The pruning pass used Step 3.7 Flash REAP observation artifacts uploaded to:
Sources:
open-r1/Mixture-of-Thoughts/{math,science,code}
and
SWE-bench/SWE-smith-trajectories/tool
Benchmark status
Terminal-Bench artifacts are uploaded separately to
0xSero/step37-prune-terminal-bench-artifacts
.
Do not treat the current Terminal-Bench evidence as a final score. The available 50B-pruned diagnostic run was interrupted and the corrected rerun hit harness/client timeouts before score-bearing proxy rows. The artifacts are useful for debugging the benchmark path, not for claiming model quality.
Loading
Use
trust_remote_code=True
and a runtime that supports Step 3.7 Flash plus ModelOpt NVFP4. For vLLM, start from StepFun's Step 3.7-compatible image and adapt the base NVFP4 launch profile:
Step-3.7-Flash-148B huggingface.co is an AI model on huggingface.co that provides Step-3.7-Flash-148B's model effect (), which can be used instantly with this 0xSero Step-3.7-Flash-148B model. huggingface.co supports a free trial of the Step-3.7-Flash-148B model, and also provides paid use of the Step-3.7-Flash-148B. Support call Step-3.7-Flash-148B model through api, including Node.js, Python, http.
Step-3.7-Flash-148B huggingface.co is an online trial and call api platform, which integrates Step-3.7-Flash-148B's modeling effects, including api services, and provides a free online trial of Step-3.7-Flash-148B, you can try Step-3.7-Flash-148B online for free by clicking the link below.
0xSero Step-3.7-Flash-148B online free url in huggingface.co:
Step-3.7-Flash-148B is an open source model from GitHub that offers a free installation service, and any user can find Step-3.7-Flash-148B on GitHub to install. At the same time, huggingface.co provides the effect of Step-3.7-Flash-148B install, users can directly use Step-3.7-Flash-148B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Step-3.7-Flash-148B install url in huggingface.co: