Quantized GGUF versions of
Qwen/Qwen2.5-Coder-7B-Instruct
— Alibaba's specialized code generation model trained on 5.5 trillion tokens of code data. Achieves performance comparable to much larger general models on coding benchmarks — the go-to for local code assistance.
Available Files
File
Quant
Size
Use Case
Qwen2.5-Coder-7B-Instruct-Q8_0.gguf
Q8_0
~7.7GB
Maximum quality
Qwen2.5-Coder-7B-Instruct-Q6_K.gguf
Q6_K
~6.0GB
Near-lossless
Qwen2.5-Coder-7B-Instruct-Q5_K_M.gguf
Q5_K_M
~5.2GB
High quality
Qwen2.5-Coder-7B-Instruct-Q4_K_M.gguf
Q4_K_M
~4.4GB
Recommended default
Qwen2.5-Coder-7B-Instruct-Q3_K_M.gguf
Q3_K_M
~3.5GB
Low VRAM
Qwen2.5-Coder-7B-Instruct-IQ4_XS.gguf
IQ4_XS
~3.9GB
Imatrix 4-bit
Qwen2.5-Coder-7B-Instruct-IQ3_XXS.gguf
IQ3_XXS
~2.9GB
Imatrix 3-bit
Qwen2.5-Coder-7B-Instruct-IQ2_M.gguf
IQ2_M
~2.5GB
Imatrix 2-bit
Qwen2.5-Coder-7B-Instruct-IQ1_S.gguf
IQ1_S
~1.8GB
Extreme compression
Qwen2.5-Coder-7B-Instruct-fp16.gguf
FP16
~14.8GB
Full precision
imatrix.dat
—
—
Importance matrix
Usage
./llama-cli -m Qwen2.5-Coder-7B-Instruct-Q4_K_M.gguf \
--ctx-size 8192 -n 1024 \
-p "<|im_start|>system\nYou are an expert programmer.<|im_end|>\n<|im_start|>user\nWrite a Python function to sort a list.<|im_end|>\n<|im_start|>assistant\n"
About Qwen2.5-Coder-7B
Parameters
: 7B
Training
: 5.5T tokens of code data
Context
: 128K tokens
License
: Apache 2.0
Strengths
: Code completion, debugging, code explanation, FIM for IDE integrations
Best-in-class local code assistance at the 7B scale.
Quantized by DuoNeural using llama.cpp on RTX 5090.
DuoNeural
DuoNeural
is an open AI research lab — human + AI in collaboration.
Qwen2.5-Coder-7B-Instruct-GGUF huggingface.co is an AI model on huggingface.co that provides Qwen2.5-Coder-7B-Instruct-GGUF's model effect (), which can be used instantly with this DuoNeural Qwen2.5-Coder-7B-Instruct-GGUF model. huggingface.co supports a free trial of the Qwen2.5-Coder-7B-Instruct-GGUF model, and also provides paid use of the Qwen2.5-Coder-7B-Instruct-GGUF. Support call Qwen2.5-Coder-7B-Instruct-GGUF model through api, including Node.js, Python, http.
Qwen2.5-Coder-7B-Instruct-GGUF huggingface.co is an online trial and call api platform, which integrates Qwen2.5-Coder-7B-Instruct-GGUF's modeling effects, including api services, and provides a free online trial of Qwen2.5-Coder-7B-Instruct-GGUF, you can try Qwen2.5-Coder-7B-Instruct-GGUF online for free by clicking the link below.
DuoNeural Qwen2.5-Coder-7B-Instruct-GGUF online free url in huggingface.co:
Qwen2.5-Coder-7B-Instruct-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Qwen2.5-Coder-7B-Instruct-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Qwen2.5-Coder-7B-Instruct-GGUF install, users can directly use Qwen2.5-Coder-7B-Instruct-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen2.5-Coder-7B-Instruct-GGUF install url in huggingface.co: