Qwen3 Coder 30B A3B
, self-quantized to GGUF by
Atomic Chat
. Built straight from Qwen's original weights with a per-tensor importance matrix. Runs fully offline.
Highlights
Agentic coding specialist
with significant performance among open models on agentic coding, agentic browser-use, and other foundational coding tasks.
Efficient MoE
: 30.5B total parameters, only 3.3B activated per token (128 experts, 8 activated).
256K native context
(262,144 tokens), extendable up to ~1M tokens with Yarn, optimized for repository-scale understanding.
Tool calling built in
with a specially designed function-call format, supporting platforms such as Qwen Code and CLINE.
Non-thinking mode only
— does not emit
<think></think>
blocks; no
enable_thinking
flag required.
Full quant ladder
with an importance matrix on every quant over
calibration_datav3
.
These GGUFs are
self-quantized from the original weights
, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
Always pass
--jinja
so the
Qwen3 Coder 30B A3B chat template
is applied. Without it the model can emit malformed turns.
Beats Q3 at similar size thanks to imatrix. Best low-RAM pick.
Q3_K_M
14.7 GB
Low quality but usable.
Q3_K_L
15.9 GB
A step above Q3_K_M.
IQ4_XS
16.4 GB
Excellent quality for size. Recommended low-bit.
Q4_K_S
17.5 GB
Compact Q4, fast.
Q4_K_M
18.6 GB
Recommended default. Best balance of size, speed and quality.
UD-Q4_K_XL
18.8 GB
Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint.
Q5_K_S
19.7 GB
Higher quality.
Q5_K_M
12.1 GB
Higher quality, low loss.
Q6_K
17.4 GB
Near lossless.
Q8_0
20.3 GB
Effectively lossless, reference quality.
Pick the largest file that fits your (V)RAM with room for context.
Q4_K_M
or
UD-Q4_K_XL
is the sweet spot for most setups;
Q6_K
or
Q8_0
for maximum fidelity.
Get started
Run Qwen3 Coder 30B A3B locally with:
Atomic Chat
:
the easiest path. Open the app, search
AtomicChat/qwen3-coder-30b-a3b-GGUF
, pick a quant, hit
Use this model
.
Qwen3-Coder-30B-A3B-GGUF huggingface.co is an AI model on huggingface.co that provides Qwen3-Coder-30B-A3B-GGUF's model effect (), which can be used instantly with this AtomicChat Qwen3-Coder-30B-A3B-GGUF model. huggingface.co supports a free trial of the Qwen3-Coder-30B-A3B-GGUF model, and also provides paid use of the Qwen3-Coder-30B-A3B-GGUF. Support call Qwen3-Coder-30B-A3B-GGUF model through api, including Node.js, Python, http.
Qwen3-Coder-30B-A3B-GGUF huggingface.co is an online trial and call api platform, which integrates Qwen3-Coder-30B-A3B-GGUF's modeling effects, including api services, and provides a free online trial of Qwen3-Coder-30B-A3B-GGUF, you can try Qwen3-Coder-30B-A3B-GGUF online for free by clicking the link below.
AtomicChat Qwen3-Coder-30B-A3B-GGUF online free url in huggingface.co:
Qwen3-Coder-30B-A3B-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-Coder-30B-A3B-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-Coder-30B-A3B-GGUF install, users can directly use Qwen3-Coder-30B-A3B-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3-Coder-30B-A3B-GGUF install url in huggingface.co: