These run on the
mobile GPU accelerator
, not the NPU.
An 8B that fits in
~2.2–2.6 GB
— ternary weights are what make a model this size practical
on a phone GPU at all.
Recommended:
bonsai-8b-int2pc-4k-gpu.litertlm
.
The plain, unmodified build with no
optimizations applied. Everything else in this repo is
experimental
— longer context or
lighter activations, but whether a given one loads depends on your app's LiteRT / LiteRT-LM
version and dependencies. Start with the recommended build; reach for an experimental one
only if you specifically need what it offers.
Ternary-Bonsai-8B's
max_position_embeddings
is
65536
, so the 32k builds here do not reach
the model's full context. A 64k build is possible and simply hasn't been made yet.
All bundles carry Bonsai's own chat template (Qwen3 ChatML with the reasoning block intact).
Why several variants
The GPU accelerator runs a
float graph
— INT2 is a
storage
format, and compute happens in
fp16/fp32. Two consequences shape this list:
Per-channel ternary dequantizes coherently.
Block-quantized weights mix scales inside a
single GEMM, which is why the per-channel builds are the conservative choice.
fp16 activations are lighter and smaller but less widely supported.
The
sdpa-fp16
bundle
is the most compact here; it is also the most likely to meet a runtime that won't take it.
Support ranges by app, so the full set is published rather than a single "best" build.
Sampling defaults
Every bundle ships these in its
LlmMetadata
, so a LiteRT-LM host picks them up without
any configuration:
parameter
value
type
TOP_P
top-k
20
top-p
0.85
temperature
0.5
These are the values the bundles were built with. Override them in your host if you want
different behaviour.
Usage
Any LiteRT-LM host — the AI Edge Gallery app, or
litert_lm_main
— with the GPU backend selected.
Built from Qwen3-8B
, Copyright 2024 Alibaba Cloud,
Apache-2.0
These bundles are quantized, repackaged derivatives. No weights were retrained.
Created using Bonsai by Prism ML.
Training data:
None was used here. These are post-training quantizations and repackagings of
the released Bonsai checkpoint; no additional training, fine-tuning, or calibration data was
involved. For the base model's training data, see the upstream Prism ML and Qwen3 model cards.
PII:
No dataset was collected, processed, or shipped as part of this conversion, so no
personally identifiable information is present in these artifacts beyond whatever the upstream
released weights already encode.
Runs of litert-community Ternary-Bonsai-8B on huggingface.co
361
Total runs
17
24-hour runs
40
3-day runs
113
7-day runs
361
30-day runs
More Information About Ternary-Bonsai-8B huggingface.co Model
Ternary-Bonsai-8B huggingface.co is an AI model on huggingface.co that provides Ternary-Bonsai-8B's model effect (), which can be used instantly with this litert-community Ternary-Bonsai-8B model. huggingface.co supports a free trial of the Ternary-Bonsai-8B model, and also provides paid use of the Ternary-Bonsai-8B. Support call Ternary-Bonsai-8B model through api, including Node.js, Python, http.
Ternary-Bonsai-8B huggingface.co is an online trial and call api platform, which integrates Ternary-Bonsai-8B's modeling effects, including api services, and provides a free online trial of Ternary-Bonsai-8B, you can try Ternary-Bonsai-8B online for free by clicking the link below.
litert-community Ternary-Bonsai-8B online free url in huggingface.co:
Ternary-Bonsai-8B is an open source model from GitHub that offers a free installation service, and any user can find Ternary-Bonsai-8B on GitHub to install. At the same time, huggingface.co provides the effect of Ternary-Bonsai-8B install, users can directly use Ternary-Bonsai-8B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.