This repository contains quantized GGUF formats of the
openbmb/BitCPM4-CANN-1B
model, heavily optimized for local inference using
llama.cpp
, text-generation-webui, LM Studio, Ollama, and other compatible backend frameworks.
The following quantization formats are available. Because this is a 1-Billion parameter model, it is highly efficient and can easily run on consumer CPUs, ultra-low-end hardware, and mobile devices.
Filename
Quant Type
File Size
Description
BitCPM4-CANN-1B-Q8_0.gguf
8-bit
1.73 GB
Highest quality, almost indistinguishable from the unquantized base model. Very fast inference.
BitCPM4-CANN-1B-Q6_K.gguf
6-bit
1.33 GB
Excellent quality with near-zero noticeable degradation. Highly recommended.
BitCPM4-CANN-1B-Q5_K_M.gguf
5-bit
1.16 GB
Great balance of file size, text generation speed, and logic retention.
BitCPM4-CANN-1B-Q5_K_S.gguf
5-bit
1.14 GB
Minor variant of
Q5_K_M
optimized slightly more for size.
BitCPM4-CANN-1B-Q4_K_M.gguf
4-bit
1.00 GB
Recommended.
The ideal sweet-spot for 4-bit formats, striking an incredible performance-to-size ratio.
BitCPM4-CANN-1B-Q4_K_S.gguf
4-bit
958 MB
Extremely small and fast. Drops below the 1GB mark, making it perfect for lightweight deployments.
BitCPM4-CANN-1B-Q3_K_M.gguf
3-bit
824 MB
Maximum compression. Use only if working under severe memory bottlenecks.
How to Run
Using
llama.cpp
(Command Line)
If you have compiled
llama.cpp
, you can run the model directly from your terminal. Replace the filename with the specific version you downloaded:
./llama-cli \
-m BitCPM4-CANN-1B-Q4_K_M.gguf \
-p "Explain the concept of artificial intelligence to a five-year-old." \
-n 256 \
-c 2048 \
--temp 0.7
Runs of Abiray BitCPM4-CANN-1B-GGUF on huggingface.co
34
Total runs
-1
24-hour runs
-3
3-day runs
-18
7-day runs
-27
30-day runs
More Information About BitCPM4-CANN-1B-GGUF huggingface.co Model
BitCPM4-CANN-1B-GGUF huggingface.co is an AI model on huggingface.co that provides BitCPM4-CANN-1B-GGUF's model effect (), which can be used instantly with this Abiray BitCPM4-CANN-1B-GGUF model. huggingface.co supports a free trial of the BitCPM4-CANN-1B-GGUF model, and also provides paid use of the BitCPM4-CANN-1B-GGUF. Support call BitCPM4-CANN-1B-GGUF model through api, including Node.js, Python, http.
BitCPM4-CANN-1B-GGUF huggingface.co is an online trial and call api platform, which integrates BitCPM4-CANN-1B-GGUF's modeling effects, including api services, and provides a free online trial of BitCPM4-CANN-1B-GGUF, you can try BitCPM4-CANN-1B-GGUF online for free by clicking the link below.
Abiray BitCPM4-CANN-1B-GGUF online free url in huggingface.co:
BitCPM4-CANN-1B-GGUF is an open source model from GitHub that offers a free installation service, and any user can find BitCPM4-CANN-1B-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of BitCPM4-CANN-1B-GGUF install, users can directly use BitCPM4-CANN-1B-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
BitCPM4-CANN-1B-GGUF install url in huggingface.co: