4-bit MLX conversion of
Qwen/Qwen3.8-Flash-Next
,
converted with
mlx-vlm
main
(post-#2032) at commit
d1bd74ed
, group size 32.
Group size 32 is required so the PLE n-gram embedding dimensions can be quantized.
Usage
pip install git+https://github.com/Blaizzy/mlx-vlm
mlx_vlm.generate --model mlx-community/Qwen3.8-Flash-Next-4bit \
--prompt "Explain sparse attention in one paragraph." --max-tokens 256
Why this conversion exists
Qwen4ExpRMSNorm
applies
1 + w
to norm gains that the released checkpoint stores
centered at zero, matching upstream
Qwen4ExpTextRMSNorm
. Several MLX conversions
published before
#2032
landed were made
with a converter that folded that
+1
into the saved weights, so loading them applies
the offset twice and generation degenerates into noise
(
#2041
).
Verification
Checked against the bf16 source after conversion:
check
result
source integrity
131/131 shards, tensor bytes byte-exact vs
index total_size
norm gain center (source vs converted)
+0.2216
vs
+0.2216
, delta
+0.00000
norm tensors bit-identical
148 / 148
sanitize()
idempotence
passes over 480 1-D tensors
generation
coherent output at
--temperature 0.0
A gain center near
+1.15
instead of
+0.22
is the signature of the double-shifted
conversions described above.
Runs of mlx-community Qwen3.8-Flash-Next-4bit on huggingface.co
7.9K
Total runs
0
24-hour runs
560
3-day runs
1.9K
7-day runs
7.9K
30-day runs
More Information About Qwen3.8-Flash-Next-4bit huggingface.co Model
Qwen3.8-Flash-Next-4bit huggingface.co
Qwen3.8-Flash-Next-4bit huggingface.co is an AI model on huggingface.co that provides Qwen3.8-Flash-Next-4bit's model effect (), which can be used instantly with this mlx-community Qwen3.8-Flash-Next-4bit model. huggingface.co supports a free trial of the Qwen3.8-Flash-Next-4bit model, and also provides paid use of the Qwen3.8-Flash-Next-4bit. Support call Qwen3.8-Flash-Next-4bit model through api, including Node.js, Python, http.
Qwen3.8-Flash-Next-4bit huggingface.co is an online trial and call api platform, which integrates Qwen3.8-Flash-Next-4bit's modeling effects, including api services, and provides a free online trial of Qwen3.8-Flash-Next-4bit, you can try Qwen3.8-Flash-Next-4bit online for free by clicking the link below.
mlx-community Qwen3.8-Flash-Next-4bit online free url in huggingface.co:
Qwen3.8-Flash-Next-4bit is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.8-Flash-Next-4bit on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.8-Flash-Next-4bit install, users can directly use Qwen3.8-Flash-Next-4bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3.8-Flash-Next-4bit install url in huggingface.co: