Embeddings, attention projections, MLP projections, LM head
License
Apache 2.0
Quantization Format: 1-bit g128
Each weight is a single bit:
0
maps to
−scale
,
1
maps to
+scale
. Every group of 128 weights shares one FP16 scale factor.
MLX's quantization formats generally store both a scale and a bias per group:
w = mlx_scale * bit + mlx_bias
. To pack our scale-only 1-bit weights into this format:
This reconstructs
−scale
when
bit=0
and
+scale
when
bit=1
. Because MLX stores two FP16 values per group (scale + bias) instead of one, the effective bits per weight is slightly higher than the GGUF format:
MLX 1-bit g128
:
1.25 bpw
(1 sign bit + two 16-bit values amortized over 128 weights)
GGUF Q1_0_g128
:
1.125 bpw
(1 sign bit + one 16-bit scale amortized over 128 weights)
Memory Requirement
Parameter memory only (weights and scales loaded into memory):
Format
Size
Reduction
Ratio
FP16
3.44 GB
—
1.0x
MLX 1-bit g128
0.27 GB
92.2%
12.8x
GGUF Q1_0_g128
0.24 GB
93.0%
14.2x
The model directory on disk is
0.28 GB (
16 MB larger) because it also includes tokenizer, config, and other metadata files alongside the weights.
Best Practices
Generation Parameters
Parameter
Default
Suggested range
Temperature
0.5
0.5 -- 0.7
Top-k
20
20 -- 40
Top-p
0.9
0.85 -- 0.95
Repetition penalty
1.0
Presence penalty
0.0
System Prompt
You can use a simple system prompt such as:
You are a helpful assistant
Quickstart
MLX (Python)
Requires PrismML fork of MLX
with 1-bit kernel support (upstream PR pending):
@techreport{bonsai,
title = {Bonsai: End-to-End 1-bit Language Model Deployment
Across Apple, GPU, and Mobile Runtimes},
author = {Prism ML},
year = {2026},
month = {March},
url = {https://prismml.com}
}
Contact
For questions, feedback, or collaboration inquiries:
[email protected]
Runs of prism-ml Bonsai-1.7B-mlx-1bit on huggingface.co
2.9K
Total runs
191
24-hour runs
697
3-day runs
749
7-day runs
706
30-day runs
More Information About Bonsai-1.7B-mlx-1bit huggingface.co Model
Bonsai-1.7B-mlx-1bit huggingface.co is an AI model on huggingface.co that provides Bonsai-1.7B-mlx-1bit's model effect (), which can be used instantly with this prism-ml Bonsai-1.7B-mlx-1bit model. huggingface.co supports a free trial of the Bonsai-1.7B-mlx-1bit model, and also provides paid use of the Bonsai-1.7B-mlx-1bit. Support call Bonsai-1.7B-mlx-1bit model through api, including Node.js, Python, http.
Bonsai-1.7B-mlx-1bit huggingface.co is an online trial and call api platform, which integrates Bonsai-1.7B-mlx-1bit's modeling effects, including api services, and provides a free online trial of Bonsai-1.7B-mlx-1bit, you can try Bonsai-1.7B-mlx-1bit online for free by clicking the link below.
prism-ml Bonsai-1.7B-mlx-1bit online free url in huggingface.co:
Bonsai-1.7B-mlx-1bit is an open source model from GitHub that offers a free installation service, and any user can find Bonsai-1.7B-mlx-1bit on GitHub to install. At the same time, huggingface.co provides the effect of Bonsai-1.7B-mlx-1bit install, users can directly use Bonsai-1.7B-mlx-1bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Bonsai-1.7B-mlx-1bit install url in huggingface.co: