All quants made using imatrix option, with a calibration corpus rendered through this model's own chat template. The corpus pairs plain prose with tool-calling and reasoning conversations (
corpus source data
), encoded exactly as this model sees them at inference and processed with
--parse-special
, so chat-format special tokens contribute to the importance matrix. The corpus rendered for this model is included in this repo:
Ling-3.0-tiny-calibration-v6.txt
. The imatrix is available here:
Ling-3.0-tiny-imatrix.gguf
.
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
ARM/AVX information
llama.cpp automatically "repacks" weights into an interleaved layout at load time for faster inference on ARM and AVX machines - details in
this PR
. This once required downloading special Q4_0_4_4/4_8/8_8 files; those are long gone. Online repacking now covers Q4_0, IQ4_NL, and most K-quants, so no special quant choice is needed for CPU inference.
Which file should I choose?
Click here for details
An older (early 2024) but still useful write-up with charts comparing quant performances is provided by Artefact2
here
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Hugging Face can also do this math for you: add your hardware in your
Local Apps settings
and the model page will show which files fit.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Ling-3.0-tiny-GGUF huggingface.co is an AI model on huggingface.co that provides Ling-3.0-tiny-GGUF's model effect (), which can be used instantly with this bartowski Ling-3.0-tiny-GGUF model. huggingface.co supports a free trial of the Ling-3.0-tiny-GGUF model, and also provides paid use of the Ling-3.0-tiny-GGUF. Support call Ling-3.0-tiny-GGUF model through api, including Node.js, Python, http.
Ling-3.0-tiny-GGUF huggingface.co is an online trial and call api platform, which integrates Ling-3.0-tiny-GGUF's modeling effects, including api services, and provides a free online trial of Ling-3.0-tiny-GGUF, you can try Ling-3.0-tiny-GGUF online for free by clicking the link below.
bartowski Ling-3.0-tiny-GGUF online free url in huggingface.co:
Ling-3.0-tiny-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Ling-3.0-tiny-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Ling-3.0-tiny-GGUF install, users can directly use Ling-3.0-tiny-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.