FP16 safetensors (HuggingFace format) of the 1-bit Bonsai 27B model. This repo exists for users who want to run Bonsai with stock HuggingFace tooling or frameworks that don't yet support 1-bit weights natively. The 1-bit hybrid-attention kernels are currently in our forks of
MLX
,
mlx-swift
, and
llama.cpp
— once they land upstream, this unpacked version will no longer be needed.
We strongly recommend using the native 1-bit models instead.
The 1-bit format is where all the benefits of Bonsai come from — a 14.2x memory reduction to 3.9 GB, interactive decoding on everyday laptops (44 tok/s on an M5 Pro), and the first 27B-class model that runs on a phone (11 tok/s on iPhone 17 Pro Max). This unpacked FP16 version is full-size (~54 GB) and does not provide any of those advantages.
For the optimized 1-bit release models (recommended):
Bonsai-27B-unpacked huggingface.co is an AI model on huggingface.co that provides Bonsai-27B-unpacked's model effect (), which can be used instantly with this prism-ml Bonsai-27B-unpacked model. huggingface.co supports a free trial of the Bonsai-27B-unpacked model, and also provides paid use of the Bonsai-27B-unpacked. Support call Bonsai-27B-unpacked model through api, including Node.js, Python, http.
Bonsai-27B-unpacked huggingface.co is an online trial and call api platform, which integrates Bonsai-27B-unpacked's modeling effects, including api services, and provides a free online trial of Bonsai-27B-unpacked, you can try Bonsai-27B-unpacked online for free by clicking the link below.
prism-ml Bonsai-27B-unpacked online free url in huggingface.co:
Bonsai-27B-unpacked is an open source model from GitHub that offers a free installation service, and any user can find Bonsai-27B-unpacked on GitHub to install. At the same time, huggingface.co provides the effect of Bonsai-27B-unpacked install, users can directly use Bonsai-27B-unpacked installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Bonsai-27B-unpacked install url in huggingface.co: