avlp12 / Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM

huggingface.co
Total runs: 805
24-hour runs: -25
7-day runs: -176
30-day runs: -59
Model's Last Updated: August 28 2026
image-text-to-text

Introduction of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM

Model Details of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM

A sensitivity-graded 3.6-bit MLX quantization of moonshotai/Kimi-K2.7-Code — a ~1T-parameter (32B active) DeepSeek-V3-style MoE model with a MoonViT vision tower — that keeps the vision encoder , so it does image-text-to-text (and video) on Apple-Silicon M3 Ultra .

465.9 GB on disk. Loads on a single clean 512 GB M3 Ultra (peak 467 GB ), or split across two over Thunderbolt.

quality

This is the vision-capable sibling of the text-only Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw . Unlike every other community MLX build of Kimi-K2.x — which drop the vision tower — this one keeps the full MoonViT encoder + multimodal projector (335 tensors, bf16) so the model can actually see.

What works
Modality Status
Text / code ✅ full (same LLM as the text build)
Image ✅ validated — correct, detailed descriptions; ~23 tok/s decode, peak 467 GB
Video ✅ the MoonViT 3D path (temporal pos-emb + spatial-temporal attention + temporal-pool merger + block-diagonal varlen attention) is ported into mlx-vlm , runs at ~23 tok/s decode, and with multi-chunk input the model reasons about motion/changes across frames.

Image example (test photo: a person from behind in a knit beanie + tan corduroy jacket, foggy forest):

"a single person photographed from behind … a thick, chunky knit beanie in muted gray … a tan/caramel-brown corduroy jacket with a prominent hood … the background is soft, blurred and misty/foggy …"

Recipe (verified from config.json )

recipe

Same sensitivity-graded LLM recipe as the text build, plus the vision tower kept in bf16 :

Component Bits
Routed experts gate / up 3-bit g64
Routed experts down_proj 4-bit on 16/60 layers, 3-bit elsewhere
Attention (MLA) · shared · dense · embed · head 6-bit g64
MoE router bf16
MoonViT vision tower + mm_projector bf16 (335 tensors, ~0.9 GB)

Effective 3.629 bits/weight . Re-quantized from the INT4 master via the #907 dequant-first fix (asking for 3-bit on a compressed-tensors source otherwise silently keeps the experts at 4-bit → ~5 bpw / 640 GB).

Usage

Needs an mlx-vlm with the kimi_k25 model (0.6.3+) and tiktoken / blobfile for the tokenizer.

Image (fast native path):

python -m mlx_vlm generate \
  --model avlp12/Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM \
  --prompt "Describe this image." --image photo.jpg \
  --max-tokens 512 --temperature 0.0

Video — mlx-vlm 0.6.3's stock kimi_k25 vision tower is 2D-only and its video CLI is Qwen-specific. This build ships a patched vision.py (3D MoonViT: temporal pos-emb, per-frame RoPE, t·h·w cu_seqlens, sd2_tpool temporal-pool merger, block-diagonal varlen attention so multi-chunk doesn't OOM) + a helper video_infer.py that decodes frames (cv2), groups them into temporal chunks, and runs the native fast generate path:

python video_infer.py <model> --video clip.mp4 --num-frames 16 --chunk-frames 4 \
       --prompt "Describe what happens over time in this video."

How it works: each chunk of ≤4 frames is temporally mean-pooled to one spatial token set ( sd2_tpool ); using multiple chunks gives the LLM a temporal sequence, so it reasons about movement/changes across frames. Decode runs at the same ~23 tok/s as images (the helper wraps generation in wired_limit to keep the 465 GB weights resident — without it decode pages from mmap and collapses to ~0.2 tok/s).

Memory

memory

MLA keeps the KV cache tiny (≈68.6 KB/token); native context 256K . Peak at image inference 467 GB on a single 512 GB box; the LLM also supports pipeline/tensor-parallel split across two machines (the vision tower is tiny and runs on one).

Credits & citation

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM — sensitivity-graded 3.6-bit image+video MLX quantization of Kimi-K2.7-Code (vision tower kept), 2026.

Runs of avlp12 Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM on huggingface.co

805
Total runs
-25
24-hour runs
-82
3-day runs
-176
7-day runs
-59
30-day runs

More Information About Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM huggingface.co Model

More Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM license Visit here:

https://choosealicense.com/licenses/modified-mit

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM huggingface.co

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM huggingface.co is an AI model on huggingface.co that provides Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM's model effect (), which can be used instantly with this avlp12 Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM model. huggingface.co supports a free trial of the Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM model, and also provides paid use of the Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM. Support call Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM model through api, including Node.js, Python, http.

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM huggingface.co Url

https://huggingface.co/avlp12/Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM

avlp12 Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM online free

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM huggingface.co is an online trial and call api platform, which integrates Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM's modeling effects, including api services, and provides a free online trial of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM, you can try Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM online for free by clicking the link below.

avlp12 Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM online free url in huggingface.co:

https://huggingface.co/avlp12/Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM install

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM is an open source model from GitHub that offers a free installation service, and any user can find Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM on GitHub to install. At the same time, huggingface.co provides the effect of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM install, users can directly use Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM install url in huggingface.co:

https://huggingface.co/avlp12/Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM

Url of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM

Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM huggingface.co Url

Provider of Kimi-K2.7-Code-Alis-MLX-Dynamic-3.6bpw-VLM huggingface.co

avlp12
ORGANIZATIONS

Other API from avlp12