philipjohnbasile / DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly

huggingface.co
Total runs: 2.9K
24-hour runs: 0
7-day runs: 2.0K
30-day runs: 2.9K
Model's Last Updated: August 14 2026
text-generation

Introduction of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly

Model Details of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly

DeepSeek V4 Flash 0731 — MLX M5 Max Target-Only (Experimental)

This is an experimental, target-only MLX derivative of deepseek-ai/DeepSeek-V4-Flash-0731 , pinned to upstream revision 7872f01b1d1fe23eabc4c98b48bffcef5a386062 .

It is intended for public testing on a 128 GB Apple Silicon Mac. It has passed real M5 Max 128 GB load-and-generation smokes, but it has not yet passed a fresh blind quality evaluation or a representative performance benchmark. Do not interpret this upload as a no-quality-loss or speed claim.

What is included
  • 43 target-model layers in 44 safetensor shards
  • 2,320 target tensors
  • 103,848,946,780 tensor-payload bytes
  • 103,855,768,335 total logical bytes in the validated local view
  • serial target model only: num_nextn_predict_layers=0
  • no MTP/DSpark drafter weights

Quantization recipe:

  • expert w1 / w3 , layers 0–38: Q2 group 128
  • expert w2 , layers 0–38: Q3 group 128
  • expert w1 / w2 / w3 , layers 39–42: Q4 group 64
  • attention, shared-expert, embedding, and head projections: affine Q8 group 64
Verified hardware smoke

Observed on an Apple M5 Max with 128 GB unified memory using the companion ReleaseFast mlx-serve DeepSeek-V4 runtime:

  • all 2,320 tensors from all 44 shards loaded
  • model reached Model ready
  • deterministic prompt output: exactly READY
  • first-touch prompt: 10 tokens at 0.937 tokens/s
  • decode: 2 tokens at 20.326 tokens/s
  • peak memory: 100.181 GB

This two-token decode is a smoke result, not a throughput benchmark. Only a 128 GB M5 Max has been tested; smaller-memory Macs are not claimed supported.

Original-model comparison

We ran the same four public, deterministic prompts against this artifact and the same DeepSeek-V4-Flash-0731 model served by OpenRouter's pinned CoreWeave FP8 endpoint. Provider fallbacks were disabled. This is a small behavioral regression gate, not a reproduction of DeepSeek's agent benchmarks and not a full-logit or source-exactness claim. The comparison was run through both local paths (direct mlx-serve and the MTPLX-routed backend) against the same frozen OpenRouter reference; see receipts/dsv4/2026-08-11-three-arm-20260812T024136Z.json .

Public case OpenRouter original Direct mlx-serve MTPLX-routed Comparison
exact single-token instruction pass pass pass byte-identical output
punctuation/case copy pass pass pass byte-identical output
constrained JSON pass pass pass parsed JSON objects equal (whitespace-only difference)
one-sentence Spanish explanation pass pass pass both valid; semantically equivalent wording

Direct mlx-serve and the MTPLX backend returned byte-identical output on all four cases (they delegate to the same engine), so the table collapses to one local column for behavior.

The original four-case remote run cost $0.00003432 and is reused here (frozen oracle, no re-spend; local arms were re-run on 2026-08-12). Timing was not sealed into this comparison receipt and network latency is never treated as model speed. The original first-party DeepSeek endpoint was unavailable under the test key's OpenRouter data-policy settings; the remote reference was the exact model slug on a pinned CoreWeave FP8 endpoint with no fallback.

Broader public-task evaluation remains pending. In particular, this artifact does not claim the upstream Terminal Bench, NL2Repo, Cybergym, DeepSWE, Toolathlon, Agents' Last Exam, AutomationBench, or DSBench scores.

Running it

The validated baseline uses the companion native mlx-serve DeepSeek-V4 runtime with prompt lookup, lossy decode-attention quantization, and vision all disabled:

mlx-serve \
  --model /path/to/DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly \
  --prompt 'Reply with exactly: READY' \
  --max-tokens 4 --temp 0 \
  --no-pld --no-decode-attn-quant --no-vision \
  --ctx-size 512 --timeout 300

Experimental MTPLX support is available via the codex/deepseek-v4-mlxserve-backend branch (backend release + gate/streaming fixes: 14413c2 ; backend initial release: 0bbe062e3a25a5de8cb31f3e8948c76516ff8404 ). It delegates to the companion native mlx-serve DeepSeek-V4 runtime and keeps this target-only artifact on the AR path. The backend sets MLX_SERVE_WIRED=fit for the child by default (override with MTPLX_DSV4_WIRED ):

mtplx pull philipjohnbasile/DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly

MTPLX_MLX_SERVE_BIN=/path/to/mlx-serve \
mtplx serve \
  --model philipjohnbasile/DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly \
  --no-mtp --host 127.0.0.1 --port 8000 --yes

Two short non-streaming MTPLX smokes completed at about 23.5 output tokens/s. An earlier streaming collapse (0.462 tokens/s then a zero-headroom refusal) was traced to the launch memory policy, not the model: with MLX_SERVE_WIRED=off the unwired 100 GB working set thrashes. Under the fit wired-residency policy ( MLX_SERVE_WIRED=fit , MLX_SERVE_WIRED_SLACK_MB=0 ), the same server streams over HTTP at 28.3 decode tokens/s with no swap activity and 100.185 GB peak memory. Use that policy when serving:

MLX_SERVE_WIRED=fit MLX_SERVE_WIRED_SLACK_MB=0 \
MTPLX_MLX_SERVE_BIN=/path/to/mlx-serve \
mtplx serve \
  --model philipjohnbasile/DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly \
  --no-mtp --host 127.0.0.1 --port 8000 --yes

Representative multi-request stress and long-context streaming still warrant a dedicated load test before any throughput claim is published.

Provenance
  • local materialization SEAL SHA-256: 50cd20ae84b6c7ebe79c27e08da89c3e419fcde5b529fbd6c47f6387bbf0e79f
  • target-only input contract SHA-256: d69c6fa36d909d0bfc964324fac00e793798088bfffc71528d96c10cb45a4b3c
  • reviewed materializer SHA-256: 55d053b4daee2556aef30ea41993acde660c759a070e2af5156c7dd5af40d275
  • materializer tests SHA-256: 83f0c3916e95c02bfd623fd9c05343d2b87e9a9d781709730dddd31b75adcb53
  • tested ReleaseFast runtime binary SHA-256: 937d2ea844e3e84d81c74daefe610ddbea0976ec7ca0de360a6e57ffb6e28201

The source model is licensed under MIT; see LICENSE and the upstream model card for attribution and its original terms.

Test feedback

When reporting a result, please include:

  • Mac model and unified-memory capacity
  • macOS version
  • runtime and model revision
  • exact flags and context length
  • whether the failure happened during load, prefill, or decode
  • peak memory and exact generated output

Runs of philipjohnbasile DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly on huggingface.co

2.9K
Total runs
0
24-hour runs
184
3-day runs
2.0K
7-day runs
2.9K
30-day runs

More Information About DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly huggingface.co Model

More DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly license Visit here:

https://choosealicense.com/licenses/mit

DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly huggingface.co

DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly huggingface.co is an AI model on huggingface.co that provides DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly's model effect (), which can be used instantly with this philipjohnbasile DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly model. huggingface.co supports a free trial of the DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly model, and also provides paid use of the DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly. Support call DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly model through api, including Node.js, Python, http.

DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly huggingface.co Url

https://huggingface.co/philipjohnbasile/DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly

philipjohnbasile DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly online free

DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly huggingface.co is an online trial and call api platform, which integrates DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly's modeling effects, including api services, and provides a free online trial of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly, you can try DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly online for free by clicking the link below.

philipjohnbasile DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly online free url in huggingface.co:

https://huggingface.co/philipjohnbasile/DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly

DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly install

DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly is an open source model from GitHub that offers a free installation service, and any user can find DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly on GitHub to install. At the same time, huggingface.co provides the effect of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly install, users can directly use DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly install url in huggingface.co:

https://huggingface.co/philipjohnbasile/DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly

Url of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly

DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly huggingface.co Url

Provider of DeepSeek-V4-Flash-0731-MLX-M5Max-TargetOnly huggingface.co

philipjohnbasile
ORGANIZATIONS

Other API from philipjohnbasile