This repository contains quantized GGUF formats of the
openbmb/BitCPM4-CANN-0.5B
model, heavily optimized for local inference using
llama.cpp
, text-generation-webui, LM Studio, Ollama, and other compatible backend frameworks.
The following quantization formats are available. Because this is an ultra-compact 500M parameter model, it can run blazingly fast on almost any modern device, including microcontrollers, older smartphones, and edge computing hardware.
Filename
Quant Type
File Size
Description
BitCPM4-CANN-0.5B-F16.gguf
16-bit
870 MB
The unquantized base model weights in full precision. Maximum possible fidelity.
BitCPM4-CANN-0.5B-Q8_0.gguf
8-bit
463 MB
Near-perfect accuracy retention. Offers a massive size reduction while acting indistinguishably from the F16 version.
BitCPM4-CANN-0.5B-Q6_K.gguf
6-bit
358 MB
Excellent option for low-resource edge devices demanding strong logic retention.
BitCPM4-CANN-0.5B-Q5_K_M.gguf
5-bit
317 MB
Great middle-ground for balancing speed, size, and remaining reasoning capability.
BitCPM4-CANN-0.5B-Q5_K_S.gguf
5-bit
310 MB
Slightly more aggressive 5-bit compression format focused on minimizing footprint.
BitCPM4-CANN-0.5B-Q4_K_M.gguf
4-bit
279 MB
Recommended.
The absolute sweet spot for local 4-bit execution, maintaining surprising coherence for its sub-300MB size.
BitCPM4-CANN-0.5B-Q4_K_S.gguf
4-bit
267 MB
Highly optimized for speed. Perfect for deeply embedded systems or background text processing.
BitCPM4-CANN-0.5B-Q3_K_M.gguf
3-bit
235 MB
Ultimate compression limit. Use exclusively under extremely severe hardware memory limitations.
How to Run
Using
llama.cpp
(Command Line)
If you have compiled
llama.cpp
, you can run the model directly from your terminal. Replace the filename with the specific version you downloaded:
./llama-cli \
-m BitCPM4-CANN-0.5B-Q4_K_M.gguf \
-p "Explain the concept of artificial intelligence to a five-year-old." \
-n 256 \
-c 2048 \
--temp 0.7
Runs of Abiray BitCPM4-CANN-0.5B-GGUF on huggingface.co
128
Total runs
-8
24-hour runs
43
3-day runs
42
7-day runs
82
30-day runs
More Information About BitCPM4-CANN-0.5B-GGUF huggingface.co Model
BitCPM4-CANN-0.5B-GGUF huggingface.co is an AI model on huggingface.co that provides BitCPM4-CANN-0.5B-GGUF's model effect (), which can be used instantly with this Abiray BitCPM4-CANN-0.5B-GGUF model. huggingface.co supports a free trial of the BitCPM4-CANN-0.5B-GGUF model, and also provides paid use of the BitCPM4-CANN-0.5B-GGUF. Support call BitCPM4-CANN-0.5B-GGUF model through api, including Node.js, Python, http.
BitCPM4-CANN-0.5B-GGUF huggingface.co is an online trial and call api platform, which integrates BitCPM4-CANN-0.5B-GGUF's modeling effects, including api services, and provides a free online trial of BitCPM4-CANN-0.5B-GGUF, you can try BitCPM4-CANN-0.5B-GGUF online for free by clicking the link below.
Abiray BitCPM4-CANN-0.5B-GGUF online free url in huggingface.co:
BitCPM4-CANN-0.5B-GGUF is an open source model from GitHub that offers a free installation service, and any user can find BitCPM4-CANN-0.5B-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of BitCPM4-CANN-0.5B-GGUF install, users can directly use BitCPM4-CANN-0.5B-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
BitCPM4-CANN-0.5B-GGUF install url in huggingface.co: