Introduction of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly
Model Details of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly
DeepSeek V4 Flash 0731 — MLX M5 Max Target-Only (Experimental)
This is an
experimental, target-only MLX derivative
of
deepseek-ai/DeepSeek-V4-Flash-0731
,
pinned to upstream revision
7872f01b1d1fe23eabc4c98b48bffcef5a386062
.
It is intended for public testing on a 128 GB Apple Silicon Mac. It has passed
real M5 Max 128 GB load-and-generation smokes, but it has
not
yet passed a
fresh blind quality evaluation or a representative performance benchmark.
Do not interpret this upload as a no-quality-loss or speed claim.
What is included
43 target-model layers in 44 safetensor shards
2,320 target tensors
103,848,946,780 tensor-payload bytes
103,855,768,335 total logical bytes in the validated local view
serial target model only:
num_nextn_predict_layers=0
no MTP/DSpark drafter weights
Quantization recipe:
expert
w1
/
w3
, layers 0–38: Q2 group 128
expert
w2
, layers 0–38: Q3 group 128
expert
w1
/
w2
/
w3
, layers 39–42: Q4 group 64
attention, shared-expert, embedding, and head projections: affine Q8 group 64
Verified hardware smoke
Observed on an Apple M5 Max with 128 GB unified memory using the companion
ReleaseFast
mlx-serve
DeepSeek-V4 runtime:
all 2,320 tensors from all 44 shards loaded
model reached
Model ready
deterministic prompt output: exactly
READY
first-touch prompt: 10 tokens at 0.937 tokens/s
decode: 2 tokens at 20.326 tokens/s
peak memory: 100.181 GB
This two-token decode is a smoke result, not a throughput benchmark. Only a
128 GB M5 Max has been tested; smaller-memory Macs are not claimed supported.
Original-model comparison
We ran the same four public, deterministic prompts against this artifact and
the same DeepSeek-V4-Flash-0731 model served by OpenRouter's pinned CoreWeave
FP8 endpoint. Provider fallbacks were disabled. This is a small behavioral
regression gate, not a reproduction of DeepSeek's agent benchmarks and not a
full-logit or source-exactness claim. The comparison was run through
both
local paths (direct
mlx-serve
and the MTPLX-routed backend) against the same
frozen OpenRouter reference; see
receipts/dsv4/2026-08-11-three-arm-20260812T024136Z.json
.
Direct
mlx-serve
and the MTPLX backend returned byte-identical output on all
four cases (they delegate to the same engine), so the table collapses to one
local column for behavior.
The original four-case remote run cost
$0.00003432
and is reused here
(frozen oracle, no re-spend; local arms were re-run on 2026-08-12). Timing was
not sealed into this comparison receipt and network latency is never treated as
model speed. The original first-party DeepSeek endpoint was unavailable under
the test key's OpenRouter data-policy settings; the remote reference was the
exact model slug on a pinned CoreWeave FP8 endpoint with no fallback.
Broader public-task evaluation remains pending. In particular, this artifact
does not claim the upstream Terminal Bench, NL2Repo, Cybergym, DeepSWE,
Toolathlon, Agents' Last Exam, AutomationBench, or DSBench scores.
Running it
The validated baseline uses the companion native
mlx-serve
DeepSeek-V4
runtime with prompt lookup, lossy decode-attention quantization, and vision all
disabled:
Experimental MTPLX support is available via the
codex/deepseek-v4-mlxserve-backend
branch (backend release + gate/streaming fixes:
14413c2
; backend initial
release:
0bbe062e3a25a5de8cb31f3e8948c76516ff8404
).
It delegates to the companion native
mlx-serve
DeepSeek-V4 runtime and keeps
this target-only artifact on the AR path. The backend sets
MLX_SERVE_WIRED=fit
for the child by default (override with
MTPLX_DSV4_WIRED
):
Two short non-streaming MTPLX smokes completed at about 23.5 output tokens/s.
An earlier streaming collapse (0.462 tokens/s then a zero-headroom refusal) was
traced to the launch memory policy, not the model: with
MLX_SERVE_WIRED=off
the unwired 100 GB working set thrashes. Under the
fit
wired-residency policy
(
MLX_SERVE_WIRED=fit
,
MLX_SERVE_WIRED_SLACK_MB=0
), the same server streams
over HTTP at 28.3 decode tokens/s with no swap activity and 100.185 GB peak
memory. Use that policy when serving:
DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly's model effect (), which can be used instantly with this philipjohnbasile DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly model. huggingface.co supports a free trial of the DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly model, and also provides paid use of the DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly. Support call DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly model through api, including Node.js, Python, http.
DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly's modeling effects, including api services, and provides a free online trial of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly, you can try DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly online for free by clicking the link below.
philipjohnbasile DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly online free url in huggingface.co:
DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly install, users can directly use DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly install url in huggingface.co: