inference-optimization / Qwen3.8-Flash-Next-MEP50

huggingface.co
Total runs: 94
24-hour runs: 0
7-day runs: -24
30-day runs: 25
Model's Last Updated: August 27 2026

Introduction of Qwen3.8-Flash-Next-MEP50

Model Details of Qwen3.8-Flash-Next-MEP50

Qwen3.8-Flash-Next - 50% Expert Pruned

50% of the MoE experts pruned by router-weight magnitude using compressed-tensors .

  • Base model: Qwen/Qwen3.8-Flash-Next
  • Sparsity: 50% of routed experts removed per layer (512 -> 256 experts)
  • Layers pruned: all 48 language model layers
  • MTP layer: retained exactly as-is
  • Shared experts: untouched
  • Vision tower: untouched (dense ViT MLP, no MoE experts)
Reproduction
from compressed_tensors.entrypoints.convert import convert_checkpoint, MagnitudeExpertPruner

convert_checkpoint(
    model_stub="Qwen/Qwen3.8-Flash-Next",
    save_directory="Qwen3.8-Flash-Next-MEP50",
    converter=MagnitudeExpertPruner.from_pretrained(
        "Qwen/Qwen3.8-Flash-Next",
        router_pattern=r"language_model\.layers\.\d+\.mlp\.gate\.weight$",
        expert_pattern=r"language_model\.layers\.\d+\.mlp\.experts\.(gate_up_proj|down_proj)$",
        sparsity=0.5,
    ),
    max_workers=8,
)

Runs of inference-optimization Qwen3.8-Flash-Next-MEP50 on huggingface.co

94
Total runs
0
24-hour runs
-5
3-day runs
-24
7-day runs
25
30-day runs

More Information About Qwen3.8-Flash-Next-MEP50 huggingface.co Model

Qwen3.8-Flash-Next-MEP50 huggingface.co

Qwen3.8-Flash-Next-MEP50 huggingface.co is an AI model on huggingface.co that provides Qwen3.8-Flash-Next-MEP50's model effect (), which can be used instantly with this inference-optimization Qwen3.8-Flash-Next-MEP50 model. huggingface.co supports a free trial of the Qwen3.8-Flash-Next-MEP50 model, and also provides paid use of the Qwen3.8-Flash-Next-MEP50. Support call Qwen3.8-Flash-Next-MEP50 model through api, including Node.js, Python, http.

inference-optimization Qwen3.8-Flash-Next-MEP50 online free

Qwen3.8-Flash-Next-MEP50 huggingface.co is an online trial and call api platform, which integrates Qwen3.8-Flash-Next-MEP50's modeling effects, including api services, and provides a free online trial of Qwen3.8-Flash-Next-MEP50, you can try Qwen3.8-Flash-Next-MEP50 online for free by clicking the link below.

inference-optimization Qwen3.8-Flash-Next-MEP50 online free url in huggingface.co:

https://huggingface.co/inference-optimization/Qwen3.8-Flash-Next-MEP50

Qwen3.8-Flash-Next-MEP50 install

Qwen3.8-Flash-Next-MEP50 is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.8-Flash-Next-MEP50 on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.8-Flash-Next-MEP50 install, users can directly use Qwen3.8-Flash-Next-MEP50 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3.8-Flash-Next-MEP50 install url in huggingface.co:

https://huggingface.co/inference-optimization/Qwen3.8-Flash-Next-MEP50

Url of Qwen3.8-Flash-Next-MEP50

Provider of Qwen3.8-Flash-Next-MEP50 huggingface.co

inference-optimization
ORGANIZATIONS

Other API from inference-optimization