¹ FP16 only fits partially on GPU's 6 GB VRAM; 1-bit fits entirely in VRAM.
Energy Efficiency
Platform
Bonsai E_tg (mWh/tok)
Baseline E_tg
Advantage
RTX 4090 (CUDA)
0.276
1.134 (FP16)
4.1x
Mac M4 Pro (Metal)
0.091
0.471 (FP16)
5.1x
Benchmarks
Evaluated with EvalScope v1.4.2 + vLLM 0.15.1 on NVIDIA H100 under identical infrastructure, generation parameters, and scoring. All models are in the 6B–9B parameter range.
Model
Company
Size
Avg
MMLU-R
MuSR
GSM8K
HE+
IFEval
BFCL
Qwen 3 8B
Alibaba
16 GB
79.3
83
55
93
82.3
84.2
81
RNJ 8B
EssentialAI
16 GB
73.1
75.5
50.4
93.7
84.2
73.8
61.1
Mistral3 8B
Mistral
16 GB
71.0
73.9
53.8
87.2
67.4
75.4
45.4
Olmo 3 7B
Allen Inst
14 GB
70.9
72
56.1
92.5
79.3
37.1
38.4
1-bit Bonsai 8B
PrismML
1.15 GB
70.5
65.7
50
88
73.8
79.8
65.7
LFM2 8B
LiquidAI
16 GB
69.6
72.7
49.5
90.1
81
82.2
62.0
Llama 3.1 8B
Meta
16 GB
67.1
72.9
51.3
87.9
75
51.5
—
GLM v6 9B
ZhipuAI
16 GB
65.7
61.9
43.2
93.4
78.7
69.3
21.9
Hermes 8B
Nous Research
16 GB
65.4
67.4
52.2
82.9
51.2
65
73.5
Trinity Nano 6B
Arcee
12 GB
61.2
68.8
52.6
81.1
54
50
62.5
Marin 8B
Stanford CRFM
16 GB
56.6
64.8
42.6
86.4
51
50
—
R1-D 7B
DeepSeek
14 GB
55.1
62.5
29.1
92.7
81.7
48.8
15.4
Despite being
1/14th the size
, 1-bit Bonsai 8B is competitive with leading full-precision 8B instruct models.
Intelligence Density
Intelligence density captures the ratio of a model's capability to its deployed size:
alpha = -ln(1 - score/100) / size_GB
Model
Size
Intelligence Density (1/GB)
1-bit Bonsai 8B
1.15 GB
1.062
Qwen 3 8B
16 GB
0.098
Llama 3.1 8B
16 GB
0.074
Mistral3 8B
16 GB
0.077
Bonsai 8B achieves
10.8x higher intelligence density
than full-precision Qwen 3 8B.
Use Cases
On-device assistants
: interactive AI on laptops and phones with low latency
Mobile deployment
: runs on a wide variety of phones due to low memory footprint
Edge robotics and autonomy
: compact deployment on devices with thermal, memory, or connectivity constraints
Cost-sensitive GPU serving
: higher throughput and lower energy per token on RTX-class and datacenter GPUs
Enterprise and private inference
: local or controlled-environment inference for data residency requirements
Limitations
No native 1-bit hardware exists yet — current gains are software-kernel optimizations on general-purpose hardware
Mobile power measurement is estimated rather than hardware-metered
The full-precision benchmark frontier continues to advance; the 1-bit methodology is architecture-agnostic and will be applied to newer bases
Citation
If you use 1-bit Bonsai 8B, please cite:
@techreport{bonsai8b,
title = {1-bit Bonsai 8B: End-to-End 1-bit Language Model Deployment
Across Apple, GPU, and Mobile Runtimes},
author = {Prism ML},
year = {2026},
month = {March},
url = {https://prismml.com}
}
Contact
For questions, feedback, or collaboration inquiries:
[email protected]
Runs of prism-ml Bonsai-8B-gguf on huggingface.co
117.4K
Total runs
0
24-hour runs
-19
3-day runs
-2.3K
7-day runs
6.1K
30-day runs
More Information About Bonsai-8B-gguf huggingface.co Model
Bonsai-8B-gguf huggingface.co is an AI model on huggingface.co that provides Bonsai-8B-gguf's model effect (), which can be used instantly with this prism-ml Bonsai-8B-gguf model. huggingface.co supports a free trial of the Bonsai-8B-gguf model, and also provides paid use of the Bonsai-8B-gguf. Support call Bonsai-8B-gguf model through api, including Node.js, Python, http.
Bonsai-8B-gguf huggingface.co is an online trial and call api platform, which integrates Bonsai-8B-gguf's modeling effects, including api services, and provides a free online trial of Bonsai-8B-gguf, you can try Bonsai-8B-gguf online for free by clicking the link below.
prism-ml Bonsai-8B-gguf online free url in huggingface.co:
Bonsai-8B-gguf is an open source model from GitHub that offers a free installation service, and any user can find Bonsai-8B-gguf on GitHub to install. At the same time, huggingface.co provides the effect of Bonsai-8B-gguf install, users can directly use Bonsai-8B-gguf installed effect in huggingface.co for debugging and trial. It also supports api for free installation.