158B total params (~37B active), 149 GB on disk, fits comfortably on a single M3/M5 Ultra (256GB+).
Requires the V4 mlx-lm port
DeepSeek-V4 is a new architecture (mHC, hash-routed MoE, sqrtsoftplus, Compressor + Indexer for compressed KV) and is
not yet in stock mlx-lm
. To use this model you need the V4 port:
from mlx_lm import load, generate
model, tokenizer = load("mlx-community/DeepSeek-V4-Flash-4bit")
out = generate(model, tokenizer, prompt="Q: What is 2+2?\nA:", max_tokens=64)
print(out)
Performance
Measured on M3 Ultra (512GB) single-node, batch=1:
Stage
tok/s
Prompt processing
6.6
Generation
20.2
Peak RAM
160 GB
Generation throughput includes the
fused Metal kernel for mHC Sinkhorn
added in PR #1189 (1.83x over the Python reference).
Source quality caveat
The bf16 source weights used for this conversion were upcasted from DeepSeek's native FP8 release rather than re-quantized directly from FP8. This stacks two quantization passes (FP8 -> BF16 -> Q4) and may produce slightly worse outputs than a direct FP8 -> Q4 conversion. A re-conversion from native FP8 is planned.
DeepSeek-V4-Flash-4bit huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4-Flash-4bit's model effect (), which can be used instantly with this mlx-community DeepSeek-V4-Flash-4bit model. huggingface.co supports a free trial of the DeepSeek-V4-Flash-4bit model, and also provides paid use of the DeepSeek-V4-Flash-4bit. Support call DeepSeek-V4-Flash-4bit model through api, including Node.js, Python, http.
DeepSeek-V4-Flash-4bit huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4-Flash-4bit's modeling effects, including api services, and provides a free online trial of DeepSeek-V4-Flash-4bit, you can try DeepSeek-V4-Flash-4bit online for free by clicking the link below.
mlx-community DeepSeek-V4-Flash-4bit online free url in huggingface.co:
DeepSeek-V4-Flash-4bit is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4-Flash-4bit on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4-Flash-4bit install, users can directly use DeepSeek-V4-Flash-4bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
DeepSeek-V4-Flash-4bit install url in huggingface.co: