Qwen3.6 27B
, self-quantized to GGUF by
Atomic Chat
. Built straight from Qwen's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
Highlights
27.8B parameters
: the weights this repo quantizes.
Context length
: 262,144 tokens (256K), as published by Qwen.
64 layers
: Dense decoder.
Modalities
: the base model handles Text, Image; this repo ships text-only quants, it carries no vision projector.
Full imatrix ladder
: every quant is calibrated with an importance matrix.
Agentic Coding:
: the model now handles frontend workflows and repository-level reasoning with greater fluency and precision.
Thinking Preservation:
: we've introduced a new option to retain reasoning context from historical messages, streamlining iterative development and reducing overhead.
These GGUFs are
self-quantized from the original weights
, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
Always pass
--jinja
so the
Qwen3.6 27B chat template
is applied. Without it the model can emit malformed turns.
Model Overview
Property
Value
Base model
Qwen/Qwen3.6-27B
Parameters
27.8B
Layers
64
Context length
262,144 tokens (256K)
Vocabulary
248,320
Modalities
Text, Image in the base model; text only in this repo, it ships no vision projector
Architecture
Dense decoder, 24 attention heads over 4 KV heads,
Qwen3_5ForConditionalGeneration
Scores are Qwen's published results for the base
Qwen/Qwen3.6-27B
, not our own measurements. Quantization preserves the large majority of this;
Q4_K_M
and up stay close to full precision.
Beats Q3 at a similar size thanks to imatrix. Best low-RAM pick.
Q3_K_M
13.3 GB
Low quality but usable.
Q3_K_L
14.3 GB
A step above Q3_K_M.
IQ4_XS
15.1 GB
Excellent quality for size. Recommended low-bit.
Q4_K_S
15.6 GB
Compact 4-bit, fast.
Q4_K_M
16.5 GB
Recommended default. Best balance of size, speed and quality.
UD-Q4_K_XL
17.5 GB
Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint.
Q5_K_S
18.7 GB
Higher quality, slightly more compact than Q5_K_M.
Q5_K_M
19.2 GB
Higher quality, low loss.
Q6_K
22.1 GB
Near lossless, noticeably lighter than Q8_0.
Q8_0
28.6 GB
Effectively lossless, reference quality.
Pick the largest file that fits your (V)RAM with room for context.
Q4_K_M
or
UD-Q4_K_XL
is the sweet spot for most setups;
Q6_K
or
Q8_0
for maximum fidelity.
Get started
Run Qwen3.6 27B locally with:
Atomic Chat
:
the easiest path. Open the app, search
AtomicChat/qwen36-27b-GGUF
, pick a quant, hit
Use this model
.
Qwen3.6-27B-GGUF huggingface.co is an AI model on huggingface.co that provides Qwen3.6-27B-GGUF's model effect (), which can be used instantly with this AtomicChat Qwen3.6-27B-GGUF model. huggingface.co supports a free trial of the Qwen3.6-27B-GGUF model, and also provides paid use of the Qwen3.6-27B-GGUF. Support call Qwen3.6-27B-GGUF model through api, including Node.js, Python, http.
Qwen3.6-27B-GGUF huggingface.co is an online trial and call api platform, which integrates Qwen3.6-27B-GGUF's modeling effects, including api services, and provides a free online trial of Qwen3.6-27B-GGUF, you can try Qwen3.6-27B-GGUF online for free by clicking the link below.
AtomicChat Qwen3.6-27B-GGUF online free url in huggingface.co:
Qwen3.6-27B-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.6-27B-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.6-27B-GGUF install, users can directly use Qwen3.6-27B-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.