infosave / Bonsai-8B_2bit_cmf

huggingface.co
Total runs: 61
24-hour runs: 2
7-day runs: 28
30-day runs: -54
Model's Last Updated: August 07 2026
text-generation

Introduction of Bonsai-8B_2bit_cmf

Model Details of Bonsai-8B_2bit_cmf

Bonsai-8B — CMF q1t (2.32 GB, runs anywhere)

Bonsai-8B — packed into a single CMF file with q1t encoding (ternary 1-bit/2-bit representation).

Because the model was trained for extreme sparsity/low-bit representation, q1t is its native representation: generation is token-for-token identical to the same model stored at higher precision. The whole model is one 2.32 GB file that memory-maps straight off disk and decodes at ~20 tok/s on an Apple M-series MacBook through a whole-token Metal graph.

What is CMF?

CMF (Cortiq Model Format) is a single-file LLM container with a small pure-Rust runtime — no Python, no torch, no C++ toolchain, no CUDA install:

  • One file carries the weights, tokenizer and chat template, and checks its own integrity.
  • mmap-first : pages load lazily on first touch; start-up is fast and RAM stays near the file size.
  • Per-tensor quantization (q8 / q4 / vbit / q1t for 1-bit/2-bit trained models).
  • O(1) attention option ( --o1 ): convert attention to a constant-memory streaming operator — no retraining, weights byte-identical.
  • Skills : one file can carry a swarm of specialists sharing a base model, with self-routing. See the skills guide .

Format spec and engine: https://github.com/infosave2007/cmf Mobile app (Flutter, runs CMF on-device): https://github.com/infosave2007/cmfmobile

Install

Prebuilt binaries (macOS arm64/x86_64, Linux, Windows): https://github.com/infosave2007/cmf/releases/latest

Or with a Rust toolchain:

cargo install cortiq-cli   # needs >= 0.5.5
Download & run
# download the model file (~2.32 GB)
huggingface-cli download infosave/Bonsai-8B_2bit_cmf bonsai-8b-q1t.cmf --local-dir .

# chat (the file carries its own chat template)
cortiq run bonsai-8b-q1t.cmf --prompt "Привет"

# Apple Silicon: run the whole-token Metal graph (~20 tok/s)
CMF_GPU=1 cortiq run bonsai-8b-q1t.cmf --prompt "Привет"

# raw completion mode (no chat template)
cortiq run bonsai-8b-q1t.cmf --prompt "The capital of France is" --raw --greedy
Run as a server

cortiq serve speaks the OpenAI API, so existing clients and SDKs work unchanged:

CMF_GPU=1 cortiq serve bonsai-8b-q1t.cmf --port 8080   # + web dashboard on /
curl localhost:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "cmf",
  "messages": [{"role": "user", "content": "Explain mmap in one sentence."}]
}'
Performance (Apple M-series MacBook)
Mode Speed
decode, CMF_GPU=1 (whole-token Metal graph) ~20.2 tok/s
decode, CPU only ~4.2 tok/s
resident memory ≈ file size (mmap)
Provenance

Converted from prism-ml/Bonsai-8B (Apache-2.0) with:

cortiq convert --model prism-ml/Bonsai-8B --quant q1t --output bonsai-8b-q1t.cmf

Runs of infosave Bonsai-8B_2bit_cmf on huggingface.co

61
Total runs
2
24-hour runs
28
3-day runs
28
7-day runs
-54
30-day runs

More Information About Bonsai-8B_2bit_cmf huggingface.co Model

More Bonsai-8B_2bit_cmf license Visit here:

https://choosealicense.com/licenses/apache-2.0

Bonsai-8B_2bit_cmf huggingface.co

Bonsai-8B_2bit_cmf huggingface.co is an AI model on huggingface.co that provides Bonsai-8B_2bit_cmf's model effect (), which can be used instantly with this infosave Bonsai-8B_2bit_cmf model. huggingface.co supports a free trial of the Bonsai-8B_2bit_cmf model, and also provides paid use of the Bonsai-8B_2bit_cmf. Support call Bonsai-8B_2bit_cmf model through api, including Node.js, Python, http.

Bonsai-8B_2bit_cmf huggingface.co Url

https://huggingface.co/infosave/Bonsai-8B_2bit_cmf

infosave Bonsai-8B_2bit_cmf online free

Bonsai-8B_2bit_cmf huggingface.co is an online trial and call api platform, which integrates Bonsai-8B_2bit_cmf's modeling effects, including api services, and provides a free online trial of Bonsai-8B_2bit_cmf, you can try Bonsai-8B_2bit_cmf online for free by clicking the link below.

infosave Bonsai-8B_2bit_cmf online free url in huggingface.co:

https://huggingface.co/infosave/Bonsai-8B_2bit_cmf

Bonsai-8B_2bit_cmf install

Bonsai-8B_2bit_cmf is an open source model from GitHub that offers a free installation service, and any user can find Bonsai-8B_2bit_cmf on GitHub to install. At the same time, huggingface.co provides the effect of Bonsai-8B_2bit_cmf install, users can directly use Bonsai-8B_2bit_cmf installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Bonsai-8B_2bit_cmf install url in huggingface.co:

https://huggingface.co/infosave/Bonsai-8B_2bit_cmf

Url of Bonsai-8B_2bit_cmf

Bonsai-8B_2bit_cmf huggingface.co Url

Provider of Bonsai-8B_2bit_cmf huggingface.co

infosave
ORGANIZATIONS

Other API from infosave

huggingface.co

Total runs: 823
Run Growth: 269
Growth Rate: 34.62%
Updated:August 28 2026
huggingface.co

Total runs: 532
Run Growth: 466
Growth Rate: 100.00%
Updated:September 23 2026
huggingface.co

Total runs: 137
Run Growth: 121
Growth Rate: 100.00%
Updated:October 01 2026
huggingface.co

Total runs: 23
Run Growth: -57
Growth Rate: -271.43%
Updated:August 19 2026
huggingface.co

Total runs: 3
Run Growth: 0
Growth Rate: 0.00%
Updated:September 24 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:August 27 2026