KAT-Coder-V2.5-Dev
is an open-weight, post-trained Mixture-of-Experts (MoE) coding agent model featuring
35B total parameters with 3B activated parameters per token
, fine-tuned on top of
Qwen3.6-35B-A3B
.
⚠️
Note:
This open-weight release contains only language-model weights and operates as a
text-only
model. Vision/multimodal components are not included.
📦 Provided GGUF Files
All quantizations in this repository were converted using
llama.cpp
and optimized using an
imatrix
(Importance Matrix) calibration file to maintain high performance at lower precision levels.
Balanced 3-bit quantization with strong reasoning retention.
KAT-Coder-V2.5-Dev-IQ4_XS.gguf
18.8 GB
Great choice for systems with ~20 GB VRAM/RAM.
KAT-Coder-V2.5-Dev-IQ4_NL.gguf
19.9 GB
Non-linear 4-bit quantization optimized via imatrix.
KAT-Coder-V2.5-Dev-Q4_K_S.gguf
20.6 GB
Small 4-bit quantization.
KAT-Coder-V2.5-Dev-Q4_K_M.gguf
21.4 GB
Recommended.
Optimal balance of speed, size, and perplexity for 24GB GPUs.
KAT-Coder-V2.5-Dev-Q5_K_S.gguf
24.2 GB
5-bit small quantization with higher fidelity.
KAT-Coder-V2.5-Dev-Q5_K_M.gguf
25.0 GB
High Quality.
Near-lossless output; suitable for 32GB+ systems.
KAT-Coder-V2.5-Dev-Q6_K.gguf
30.1 GB
High-precision 6-bit quant for power users.
KAT-Coder-V2.5-Dev-Q8_0.gguf
36.9 GB
Virtually identical to full 16-bit float precision.
🚀 Quickstart & Usage
Running with
llama.cpp
Ensure you are using a recent build of
llama.cpp
that supports Qwen3 / MoE architectures.
CLI Example:
./llama-cli -m KAT-Coder-V2.5-Dev-Q4_K_M.gguf \
-p "Write a Python function that implements a binary search tree with deletion." \
-n 4096 \
-c 32768 \
--temp 0.7
Download your preferred
.gguf
file from the table above.
Place the file inside your local model folder (e.g.,
~/.cache/lm-studio/models
or KoboldCpp directory).
Set your context size up to
262,144 tokens
(adjust depending on your available system RAM/VRAM).
✨ Original Model Highlights
SOTA Agentic Coding Performance:
Through post-training SFT and RL, KAT-Coder-V2.5-Dev achieves state-of-the-art results among models of similar parameter scales on benchmark tasks like
SWE-bench Verified (69.40%)
.
Reduced Pathological Behaviors:
Reinforcement Learning significantly reduced unwanted behaviors, such as abnormal tool labels (-9pp improvement) and single-turn continuous repetitions (reduced to 0%).
Preserve Thinking Mode:
The model supports retaining historical thinking context across multi-turn interactions, improving agent consistency and saving redundant reasoning tokens.
📊 Benchmark Performance
The table below shows the official benchmark evaluation results reproduced in-house by the original authors:
Benchmark
KAT-Coder-V2.5-Dev
Qwen3.5-27B
Qwen3.6-35BA3B
Gemma4-31B
Qwen3.5-35BA3B
Ornith-1.0-35B
Gemma4-26BA4B
Qwen3-Coder-30B
SWE-bench Verified
69.40
68.60
64.40
60.60
58.60
55.80
35.80
31.80
SWE-bench Multilingual
63.00
57.67
57.00
49.33
47.67
51.67
27.33
20.67
SWE-bench Pro
45.96
42.13
40.63
32.97
38.03
34.47
9.58
19.84
Terminal-Bench 2.1
41.02
34.84
32.02
32.59
26.12
35.98
20.94
13.50
PinchBench
93.43
90.71
92.21
85.53
88.75
91.62
82.01
72.30
Scicode
44.20
25.58
37.53
33.19
27.73
30.34
30.84
18.27
KAT-Code-Bench
46.21
44.83
42.76
37.93
35.86
33.10
22.06
15.17
🔬 Post-Training Details
KAT-Coder-V2.5-Dev follows a two-stage post-training pipeline built on top of
Qwen3.6-35B-A3B
:
Supervised Fine-Tuning (SFT):
Fine-tuned on 127K curated agentic and coding examples.
Reinforcement Learning (RL):
Token-in-Token-out (TITO) Consistency:
Eliminates off-policy training discrepancies caused by tokenizer or chat-template changes.
KAT-Coder-V2.5-Dev-Imatrix-GGUF huggingface.co is an AI model on huggingface.co that provides KAT-Coder-V2.5-Dev-Imatrix-GGUF's model effect (), which can be used instantly with this Abiray KAT-Coder-V2.5-Dev-Imatrix-GGUF model. huggingface.co supports a free trial of the KAT-Coder-V2.5-Dev-Imatrix-GGUF model, and also provides paid use of the KAT-Coder-V2.5-Dev-Imatrix-GGUF. Support call KAT-Coder-V2.5-Dev-Imatrix-GGUF model through api, including Node.js, Python, http.
KAT-Coder-V2.5-Dev-Imatrix-GGUF huggingface.co is an online trial and call api platform, which integrates KAT-Coder-V2.5-Dev-Imatrix-GGUF's modeling effects, including api services, and provides a free online trial of KAT-Coder-V2.5-Dev-Imatrix-GGUF, you can try KAT-Coder-V2.5-Dev-Imatrix-GGUF online for free by clicking the link below.
Abiray KAT-Coder-V2.5-Dev-Imatrix-GGUF online free url in huggingface.co:
KAT-Coder-V2.5-Dev-Imatrix-GGUF is an open source model from GitHub that offers a free installation service, and any user can find KAT-Coder-V2.5-Dev-Imatrix-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of KAT-Coder-V2.5-Dev-Imatrix-GGUF install, users can directly use KAT-Coder-V2.5-Dev-Imatrix-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
KAT-Coder-V2.5-Dev-Imatrix-GGUF install url in huggingface.co: