Bonsai-8B
— packed into a single
CMF
file with
q1t
encoding (ternary 1-bit/2-bit representation).
Because the model was
trained
for extreme sparsity/low-bit representation, q1t is its native
representation: generation is
token-for-token identical
to the same model stored at higher precision. The whole
model is one 2.32 GB file that memory-maps straight off disk and
decodes at
~20 tok/s on an Apple M-series MacBook
through a
whole-token Metal graph.
What is CMF?
CMF (Cortiq Model Format) is a single-file LLM container with a small
pure-Rust runtime — no Python, no torch, no C++ toolchain, no CUDA
install:
One file
carries the weights, tokenizer and chat template, and
checks its own integrity.
mmap-first
: pages load lazily on first touch; start-up is fast
and RAM stays near the file size.
# download the model file (~2.32 GB)
huggingface-cli download infosave/Bonsai-8B_2bit_cmf bonsai-8b-q1t.cmf --local-dir .
# chat (the file carries its own chat template)
cortiq run bonsai-8b-q1t.cmf --prompt "Привет"# Apple Silicon: run the whole-token Metal graph (~20 tok/s)
CMF_GPU=1 cortiq run bonsai-8b-q1t.cmf --prompt "Привет"# raw completion mode (no chat template)
cortiq run bonsai-8b-q1t.cmf --prompt "The capital of France is" --raw --greedy
Run as a server
cortiq serve
speaks the OpenAI API, so existing clients and SDKs work
unchanged:
CMF_GPU=1 cortiq serve bonsai-8b-q1t.cmf --port 8080 # + web dashboard on /
curl localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{ "model": "cmf", "messages": [{"role": "user", "content": "Explain mmap in one sentence."}]}'
Performance (Apple M-series MacBook)
Mode
Speed
decode,
CMF_GPU=1
(whole-token Metal graph)
~20.2 tok/s
decode, CPU only
~4.2 tok/s
resident memory
≈ file size (mmap)
Provenance
Converted from prism-ml/Bonsai-8B (Apache-2.0) with:
Bonsai-8B_2bit_cmf huggingface.co is an AI model on huggingface.co that provides Bonsai-8B_2bit_cmf's model effect (), which can be used instantly with this infosave Bonsai-8B_2bit_cmf model. huggingface.co supports a free trial of the Bonsai-8B_2bit_cmf model, and also provides paid use of the Bonsai-8B_2bit_cmf. Support call Bonsai-8B_2bit_cmf model through api, including Node.js, Python, http.
Bonsai-8B_2bit_cmf huggingface.co is an online trial and call api platform, which integrates Bonsai-8B_2bit_cmf's modeling effects, including api services, and provides a free online trial of Bonsai-8B_2bit_cmf, you can try Bonsai-8B_2bit_cmf online for free by clicking the link below.
infosave Bonsai-8B_2bit_cmf online free url in huggingface.co:
Bonsai-8B_2bit_cmf is an open source model from GitHub that offers a free installation service, and any user can find Bonsai-8B_2bit_cmf on GitHub to install. At the same time, huggingface.co provides the effect of Bonsai-8B_2bit_cmf install, users can directly use Bonsai-8B_2bit_cmf installed effect in huggingface.co for debugging and trial. It also supports api for free installation.