Mixed-precision quantized with OptiQ — sensitivity-driven quantization for Apple Silicon
This is a mixed-precision quantized version of
google/gemma-4-e4b-it
in MLX format. OptiQ measures each layer's sensitivity via KL divergence and assigns optimal per-layer bit-widths, preserving model quality where it matters most.
Quantization Details
Property
Value
Target BPW
4.5
Achieved BPW
4.50
Candidate bits
4, 8
Layers at 4-bit
444
Layers at 8-bit
149
Total quantized layers
593
Group size
64
Benchmark Results
GSM8K
(200 samples, 3-shot chain-of-thought):
Model
GSM8K Accuracy
OptiQ mixed (4.5 BPW)
55.5%
Uniform 4-bit
23.5%
OptiQ more than doubles the accuracy of uniform 4-bit quantization (+32.0 percentage points, 2.4x improvement).
gemma-4-e4b-it-OptiQ-4bit huggingface.co is an AI model on huggingface.co that provides gemma-4-e4b-it-OptiQ-4bit's model effect (), which can be used instantly with this mlx-community gemma-4-e4b-it-OptiQ-4bit model. huggingface.co supports a free trial of the gemma-4-e4b-it-OptiQ-4bit model, and also provides paid use of the gemma-4-e4b-it-OptiQ-4bit. Support call gemma-4-e4b-it-OptiQ-4bit model through api, including Node.js, Python, http.
gemma-4-e4b-it-OptiQ-4bit huggingface.co is an online trial and call api platform, which integrates gemma-4-e4b-it-OptiQ-4bit's modeling effects, including api services, and provides a free online trial of gemma-4-e4b-it-OptiQ-4bit, you can try gemma-4-e4b-it-OptiQ-4bit online for free by clicking the link below.
mlx-community gemma-4-e4b-it-OptiQ-4bit online free url in huggingface.co:
gemma-4-e4b-it-OptiQ-4bit is an open source model from GitHub that offers a free installation service, and any user can find gemma-4-e4b-it-OptiQ-4bit on GitHub to install. At the same time, huggingface.co provides the effect of gemma-4-e4b-it-OptiQ-4bit install, users can directly use gemma-4-e4b-it-OptiQ-4bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
gemma-4-e4b-it-OptiQ-4bit install url in huggingface.co: