314.8B total parameters (13.2B active per token).
The sidebar "params" chip is currently
miscounted by the Hub for quantized MLX repos (it counts packed uint32 elements; the correct
total_parameters
is declared in
model.safetensors.index.json
— reported upstream).
8-bit (group-size 64, affine) MLX quantization of
Motif-3
— the final release of
Motif Technologies' Korean-emphasized reasoning MoE (MIT license).
262,144 (YaRN ×64 from 4,096;
apply_yarn_scaling: false
per vendor config)
Vocab
220,160 · reasoning chat template (auto-opens a
<think>
channel)
Size on disk
312 GiB (8.50 bpw)
— fits a single 512 GB M3 Ultra with ~180 GB headroom
Measured performance (M3 Ultra 512 GB, greedy)
Metric
Value
Decode
23.3–24.3 tok/s
with MTP self-spec (k=1; long-form is faster) · 20.2 plain · 14.4 eager
Prefill
~143 tok/s (834-token prompt → 5.8 s to first token)
Peak memory
~334 GB
The mHC gates' 20-iteration Sinkhorn loop is ~4,200 tiny GPU dispatches per token in the
eager path; the fork fuses it into one Metal kernel (+40% decode). Numerics vs eager:
KL 2.2e-3, 0% top-1 flips over 256 positions; kill switch
MOTIF_SINKHORN_KERNEL=0
.
Speed-tier sibling
Need more speed and can spare a little fidelity? The
4.5bpw sibling build
(167 GiB)
decodes at ~30 tok/s — about +25% over this build — at a small measured quality cost
(KO long-form NLL +2.4% vs this 8-bit reference). This 8-bit build remains the fidelity
reference.
MTP self-speculative decoding (opt-in)
This build now ships the vendor-trained MTP (nextn) block, quantized 8-bit
(
model-mtp.safetensors
, +267 MB). With the fork,
--mtp --mtp-num-draft-tokens 1
drafts one
token per step from the MTP head and verifies it in the same forward — distribution-lossless:
+15–20% decode (20.2 → 23.3 short / 24.3 long-form tok/s), ~65–70% draft acceptance (greedy) —
acceptance rises with generation length
. (An earlier revision of this card quoted "38–41%";
that figure was the fraction of
emitted
tokens coming from the draft, i.e. a/(1+a) — the
per-draft acceptance rate is ~65–70%, consistent with the 70–80% Motif reports for their own
vLLM stack.) At serving temperatures the fork additionally supports
rejection-sampling
acceptance
(Leviathan accept prob min(1, p/q), residual resample on reject — exactly
preserves the sampling distribution): on this build at T=0.8 it lifts acceptance ~60% → ~73%
and decode 22.7 → 24.5 tok/s (3-prompt avg) — i.e. temperature serving now matches greedy
speed. The server auto-enables it for pure-temperature requests; kill switch
MLX_MTP_REJECTION=0
(top-p/top-k requests fall back to equality acceptance).
k≥2 chaining is net-negative and not recommended (matches Motif's own "1 speculative token
is optimal" guidance).
The wiring is now vendor-faithful, verified against Motif's own training reference
(
MotifTechnologies/motif3-training-example
): post-final-norm hidden anchor,
[h ; embed_norm(emb)]
concat, SWA attention (window 129) for the MTP block. Note: greedy
transcripts under MTP can diverge from the plain path after many tokens (multi-token verify
kernels vs single-step kernels — ULP-class tie flips); the sampling distribution is equivalent.
All expert/attention/dense projections 8-bit g64 affine (632 tensors, per-tensor map in
config.json
). Kept in bf16: router gate,
mhc_*
,
lambda_proj
, norms, PolyNorm
coefficients. The MTP head is dropped (not instantiated by the modeling code). Tensors with
2³¹ elements are sanitized without
mx.split
(upstream mlx#3836 silent-corruption
workaround; byte-verified).
Usage notes
Sampling
: use temp ≈ 0.6–0.7. At temp 0 the think channel can enter repetition loops.
Verbatim recall
(anthems, poems, legal text) is unreliable — the model may blend
historical variants; pair with retrieval/web grounding for exact quotations.
The chat template opens
<think>
automatically; servers that split reasoning
(e.g., this fork) return it in
message.reasoning
.
Provenance & verification
Port lineage: the Motif-3-Beta port (avlp12/Motif-3-Beta-Alis-MLX-*) — 4-layer parity vs the
vendor's fixed reference at ~1e-7 KL/token, plus the PolyNorm/BUG-5 saga resolved with the
Motif team (HF discussions). The final release confirms the Beta RoPE hypothesis B
(
apply_yarn_scaling: false
, amplitude 1.0). Methodology and receipts:
alis-dwq
.
Smoke-verified on Korean and English prompts; served as a web chat (Open WebUI over
mlx_lm.server
) on a single M3 Ultra.
Runs of avlp12 Motif-3-Alis-MLX-8bit on huggingface.co
64
Total runs
1
24-hour runs
-8
3-day runs
-16
7-day runs
-240
30-day runs
More Information About Motif-3-Alis-MLX-8bit huggingface.co Model
Motif-3-Alis-MLX-8bit huggingface.co is an AI model on huggingface.co that provides Motif-3-Alis-MLX-8bit's model effect (), which can be used instantly with this avlp12 Motif-3-Alis-MLX-8bit model. huggingface.co supports a free trial of the Motif-3-Alis-MLX-8bit model, and also provides paid use of the Motif-3-Alis-MLX-8bit. Support call Motif-3-Alis-MLX-8bit model through api, including Node.js, Python, http.
Motif-3-Alis-MLX-8bit huggingface.co is an online trial and call api platform, which integrates Motif-3-Alis-MLX-8bit's modeling effects, including api services, and provides a free online trial of Motif-3-Alis-MLX-8bit, you can try Motif-3-Alis-MLX-8bit online for free by clicking the link below.
avlp12 Motif-3-Alis-MLX-8bit online free url in huggingface.co:
Motif-3-Alis-MLX-8bit is an open source model from GitHub that offers a free installation service, and any user can find Motif-3-Alis-MLX-8bit on GitHub to install. At the same time, huggingface.co provides the effect of Motif-3-Alis-MLX-8bit install, users can directly use Motif-3-Alis-MLX-8bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Motif-3-Alis-MLX-8bit install url in huggingface.co: