sakamakismile / KAT-Coder-V2.5-Dev-NVFP4

huggingface.co
Total runs: 1.0K
24-hour runs: 0
7-day runs: -781
30-day runs: -3.2K
Model's Last Updated: July 24 2026

Introduction of KAT-Coder-V2.5-Dev-NVFP4

Model Details of KAT-Coder-V2.5-Dev-NVFP4

KAT-Coder-V2.5-Dev-NVFP4

NVFP4 (W4A4, compressed-tensors) quantization of Kwaipilot/KAT-Coder-V2.5-Dev — the 35B-A3B agentic-coding MoE (SWE-bench Verified 69.4% upstream).

70GB bf16 → 21.9GB. Runs on a single 24GB Blackwell card, or 2× 16GB.

Quantized by Lna-Lab ( @Tono_Ken3 ).

Measured (12x RTX PRO 2000 Blackwell 16GB, 60W power cap each)
  • TP=2 (2 GPUs): ~122 tok/s single-stream
  • Sanity: lookahead-bias trading question answered with the correct shift(1) fix; clean O(n) implementations with tests.
Serving
vllm serve <this-repo> --tensor-parallel-size 2 --max-model-len 32768 \
  --gpu-memory-utilization 0.92
# NVFP4 is auto-detected; no --quantization flag needed.
# On no-P2P multi-GPU boxes add: NCCL_P2P_DISABLE=1, --disable-custom-all-reduce
Recipe notes
  • Arch qwen3_5_moe quantizes cleanly with llm-compressor ( targets=Linear , scheme NVFP4) with these ignores: lm_head , re:.*conv1d.* (DeltaNet conv), re:.*mlp.gate$ and re:.*shared_expert_gate$ (MoE routers), re:.*mtp.* .
  • The open-weight release ships without vision and without MTP tensors (we checked; nothing to graft).
  • Calibration: 32 samples x 8192 tokens (neuralmagic/calibration), single GPU, ~4 minutes total wall.
  • Newer transformers removed GraniteMoeParallelExperts which llm-compressor still imports — a dummy class injected before import satisfies it safely for non-Granite models.

W4A4 is stable on this architecture (5th model family we've confirmed: Nex-N2, Holo, Ornith, AgentWorld, KAT).

Runs of sakamakismile KAT-Coder-V2.5-Dev-NVFP4 on huggingface.co

1.0K
Total runs
0
24-hour runs
-145
3-day runs
-781
7-day runs
-3.2K
30-day runs

More Information About KAT-Coder-V2.5-Dev-NVFP4 huggingface.co Model

More KAT-Coder-V2.5-Dev-NVFP4 license Visit here:

https://choosealicense.com/licenses/apache-2.0

KAT-Coder-V2.5-Dev-NVFP4 huggingface.co

KAT-Coder-V2.5-Dev-NVFP4 huggingface.co is an AI model on huggingface.co that provides KAT-Coder-V2.5-Dev-NVFP4's model effect (), which can be used instantly with this sakamakismile KAT-Coder-V2.5-Dev-NVFP4 model. huggingface.co supports a free trial of the KAT-Coder-V2.5-Dev-NVFP4 model, and also provides paid use of the KAT-Coder-V2.5-Dev-NVFP4. Support call KAT-Coder-V2.5-Dev-NVFP4 model through api, including Node.js, Python, http.

KAT-Coder-V2.5-Dev-NVFP4 huggingface.co Url

https://huggingface.co/sakamakismile/KAT-Coder-V2.5-Dev-NVFP4

sakamakismile KAT-Coder-V2.5-Dev-NVFP4 online free

KAT-Coder-V2.5-Dev-NVFP4 huggingface.co is an online trial and call api platform, which integrates KAT-Coder-V2.5-Dev-NVFP4's modeling effects, including api services, and provides a free online trial of KAT-Coder-V2.5-Dev-NVFP4, you can try KAT-Coder-V2.5-Dev-NVFP4 online for free by clicking the link below.

sakamakismile KAT-Coder-V2.5-Dev-NVFP4 online free url in huggingface.co:

https://huggingface.co/sakamakismile/KAT-Coder-V2.5-Dev-NVFP4

KAT-Coder-V2.5-Dev-NVFP4 install

KAT-Coder-V2.5-Dev-NVFP4 is an open source model from GitHub that offers a free installation service, and any user can find KAT-Coder-V2.5-Dev-NVFP4 on GitHub to install. At the same time, huggingface.co provides the effect of KAT-Coder-V2.5-Dev-NVFP4 install, users can directly use KAT-Coder-V2.5-Dev-NVFP4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

KAT-Coder-V2.5-Dev-NVFP4 install url in huggingface.co:

https://huggingface.co/sakamakismile/KAT-Coder-V2.5-Dev-NVFP4

Url of KAT-Coder-V2.5-Dev-NVFP4

KAT-Coder-V2.5-Dev-NVFP4 huggingface.co Url

Provider of KAT-Coder-V2.5-Dev-NVFP4 huggingface.co

sakamakismile
ORGANIZATIONS

Other API from sakamakismile