llama.cpp downloads the mmproj automatically when using
-hf
as shown above; if you're loading files manually, pass it with
--mmproj
.
imatrix
All quants made using imatrix option, with a calibration corpus rendered through this model's own chat template. The corpus pairs plain prose with tool-calling and reasoning conversations (
corpus source data
), encoded exactly as this model sees them at inference and processed with
--parse-special
, so chat-format special tokens contribute to the importance matrix. The corpus rendered for this model is included in this repo:
Ornith-1.5-9B-calibration-v6.txt
. The imatrix is available here:
Ornith-1.5-9B-imatrix.gguf
.
Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.
ARM/AVX information
llama.cpp automatically "repacks" weights into an interleaved layout at load time for faster inference on ARM and AVX machines - details in
this PR
. This once required downloading special Q4_0_4_4/4_8/8_8 files; those are long gone. Online repacking now covers Q4_0, IQ4_NL, and most K-quants, so no special quant choice is needed for CPU inference.
Which file should I choose?
Click here for details
An older (early 2024) but still useful write-up with charts comparing quant performances is provided by Artefact2
here
The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.
If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.
If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.
Hugging Face can also do this math for you: add your hardware in your
Local Apps settings
and the model page will show which files fit.
Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.
If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.
If you want to get more into the weeds, you can check out this extremely useful feature chart:
But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.
These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.
Credits
Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.
Thank you ZeroWw for the inspiration to experiment with embed/output.
Ornith-1.5-9B-GGUF huggingface.co is an AI model on huggingface.co that provides Ornith-1.5-9B-GGUF's model effect (), which can be used instantly with this bartowski Ornith-1.5-9B-GGUF model. huggingface.co supports a free trial of the Ornith-1.5-9B-GGUF model, and also provides paid use of the Ornith-1.5-9B-GGUF. Support call Ornith-1.5-9B-GGUF model through api, including Node.js, Python, http.
Ornith-1.5-9B-GGUF huggingface.co is an online trial and call api platform, which integrates Ornith-1.5-9B-GGUF's modeling effects, including api services, and provides a free online trial of Ornith-1.5-9B-GGUF, you can try Ornith-1.5-9B-GGUF online for free by clicking the link below.
bartowski Ornith-1.5-9B-GGUF online free url in huggingface.co:
Ornith-1.5-9B-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Ornith-1.5-9B-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Ornith-1.5-9B-GGUF install, users can directly use Ornith-1.5-9B-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.