Introduction of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM
Model Details of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM
Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM
A
sensitivity-graded 3.6-bit MLX quantization
of
moonshotai/Kimi-K2.7-Code
— a ~1T-parameter (32B active) DeepSeek-V3-style MoE model with a
MoonViT
vision tower — that
keeps the vision encoder
, so it does
image-text-to-text
(and video) on Apple-Silicon
M3 Ultra
.
465.9 GB
on disk. Loads on a single clean
512 GB
M3 Ultra (peak
467 GB
), or split across two over Thunderbolt.
This is the
vision-capable
sibling of the text-only
Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw
. Unlike every other community MLX build of Kimi-K2.x — which
drop
the vision tower — this one keeps the full
MoonViT
encoder + multimodal projector (335 tensors, bf16) so the model can actually see.
✅ the MoonViT
3D
path (temporal pos-emb + spatial-temporal attention + temporal-pool merger + block-diagonal varlen attention) is ported into
mlx-vlm
, runs at
~23 tok/s
decode, and with
multi-chunk
input the model reasons about motion/changes across frames.
Image example (test photo: a person from behind in a knit beanie + tan corduroy jacket, foggy forest):
"a single person photographed from behind … a
thick, chunky knit beanie
in muted gray … a
tan/caramel-brown corduroy jacket
with a prominent hood … the background is soft, blurred and
misty/foggy
…"
Recipe (verified from
config.json
)
Same sensitivity-graded LLM recipe as the text build,
plus the vision tower kept in bf16
:
Component
Bits
Routed experts
gate
/
up
3-bit g64
Routed experts
down_proj
4-bit on 16/60 layers, 3-bit elsewhere
Attention (MLA) · shared · dense · embed · head
6-bit g64
MoE router
bf16
MoonViT vision tower + mm_projector
bf16
(335 tensors, ~0.9 GB)
Effective
3.629 bits/weight
. Re-quantized from the INT4 master via the
#907
dequant-first fix (asking for 3-bit on a
compressed-tensors
source otherwise silently keeps the experts at 4-bit → ~5 bpw / 640 GB).
Usage
Needs an
mlx-vlm
with the
kimi_k25
model (0.6.3+) and
tiktoken
/
blobfile
for the tokenizer.
Video
— mlx-vlm 0.6.3's stock
kimi_k25
vision tower is
2D-only
and its video CLI is Qwen-specific. This build ships a patched
vision.py
(3D MoonViT: temporal pos-emb, per-frame RoPE,
t·h·w
cu_seqlens,
sd2_tpool
temporal-pool merger,
block-diagonal varlen attention
so multi-chunk doesn't OOM) + a helper
video_infer.py
that decodes frames (cv2), groups them into temporal chunks, and runs the
native fast generate
path:
python video_infer.py <model> --video clip.mp4 --num-frames 16 --chunk-frames 4 \
--prompt "Describe what happens over time in this video."
How it works: each chunk of ≤4 frames is temporally mean-pooled to one spatial token set (
sd2_tpool
); using
multiple chunks
gives the LLM a temporal sequence, so it reasons about movement/changes across frames. Decode runs at the same
~23 tok/s
as images (the helper wraps generation in
wired_limit
to keep the 465 GB weights resident — without it decode pages from mmap and collapses to ~0.2 tok/s).
Memory
MLA keeps the KV cache tiny (≈68.6 KB/token); native context
256K
. Peak at image inference
467 GB
on a single 512 GB box; the LLM also supports pipeline/tensor-parallel split across two machines (the vision tower is tiny and runs on one).
Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM huggingface.co is an AI model on huggingface.co that provides Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM's model effect (), which can be used instantly with this avlp12 Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM model. huggingface.co supports a free trial of the Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM model, and also provides paid use of the Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM. Support call Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM model through api, including Node.js, Python, http.
Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM huggingface.co is an online trial and call api platform, which integrates Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM's modeling effects, including api services, and provides a free online trial of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM, you can try Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM online for free by clicking the link below.
avlp12 Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM online free url in huggingface.co:
Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM is an open source model from GitHub that offers a free installation service, and any user can find Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM on GitHub to install. At the same time, huggingface.co provides the effect of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM install, users can directly use Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM install url in huggingface.co: