PRISM Dynamic Quantization (PRISM-DQ)
applies per-tensor-class bit allocation based on structural weight analysis — no calibration data or importance matrices required. Each tensor class (attention keys, FFN gates, SSM components, etc.) receives a quantization type proportional to its measured sensitivity, while staying within a target bits-per-weight budget.
This repo contains PRISM-DQ quantized GGUFs for the full
Qwen3.5
vision-language model family (0.8B, 2B, 4B, 9B), plus multimodal projection weights (mmproj) for vision capabilities.
# Download a model
huggingface-cli download Ex0bit/Qwen3.5-PRISM-Dynamic-Quant-GGUF \
Qwen3.5-9B/Qwen3.5-9B-PRISM-DQ.gguf --local-dir .
# Run with llama-cli
llama-cli -m Qwen3.5-9B/Qwen3.5-9B-PRISM-DQ.gguf \
-p "You are a helpful assistant." \
--chat-template-file Qwen3.5-9B/chat_template.jinja \
-cnv
Vision (multimodal)
# Download model + mmproj
huggingface-cli download Ex0bit/Qwen3.5-PRISM-Dynamic-Quant-GGUF \
Qwen3.5-9B/Qwen3.5-9B-PRISM-DQ.gguf \
Qwen3.5-9B/mmproj-BF16.gguf --local-dir .
# Run with llama-mtmd-cli
llama-mtmd-cli -m Qwen3.5-9B/Qwen3.5-9B-PRISM-DQ.gguf \
--mmproj Qwen3.5-9B/mmproj-BF16.gguf \
--chat-template-file Qwen3.5-9B/chat_template.jinja \
-cnv
LM Studio / Ollama
These GGUFs work with any llama.cpp-compatible runtime. Simply point your application at the
.gguf
file.
PRISM Dynamic Quantization analyzes each weight tensor using 7 structural metrics:
PL-Alpha-Hill
— spectral heavy-tail index via eigenvalue analysis
Spectral Dominance
— top singular value ratio (rank-1 approximation quality)
OSQE
— optimal scale quantization error at multiple bit levels (2, 3, 4, 6 bit)
Matrix Imbalance
— max of row/column coefficient of variation
Fragility
— log-ratio of 2-bit vs 4-bit quantization error
Boundary Density
— fraction of values near quantization bin boundaries
Spectral Position Prior
— bidirectional spectral norm product encoding layer position
These metrics are combined into a composite sensitivity score per tensor class. A Lagrangian allocator then distributes bits across classes to minimize total quantization distortion subject to the BPW budget, with per-block refinement for individual tensor overrides.
License
This model is released under the Apache 2.0 license, consistent with the base Qwen3.5 models.
Qwen3.5-PRISM-Dynamic-Quant-GGUF huggingface.co is an AI model on huggingface.co that provides Qwen3.5-PRISM-Dynamic-Quant-GGUF's model effect (), which can be used instantly with this Ex0bit Qwen3.5-PRISM-Dynamic-Quant-GGUF model. huggingface.co supports a free trial of the Qwen3.5-PRISM-Dynamic-Quant-GGUF model, and also provides paid use of the Qwen3.5-PRISM-Dynamic-Quant-GGUF. Support call Qwen3.5-PRISM-Dynamic-Quant-GGUF model through api, including Node.js, Python, http.
Qwen3.5-PRISM-Dynamic-Quant-GGUF huggingface.co is an online trial and call api platform, which integrates Qwen3.5-PRISM-Dynamic-Quant-GGUF's modeling effects, including api services, and provides a free online trial of Qwen3.5-PRISM-Dynamic-Quant-GGUF, you can try Qwen3.5-PRISM-Dynamic-Quant-GGUF online for free by clicking the link below.
Ex0bit Qwen3.5-PRISM-Dynamic-Quant-GGUF online free url in huggingface.co:
Qwen3.5-PRISM-Dynamic-Quant-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Qwen3.5-PRISM-Dynamic-Quant-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3.5-PRISM-Dynamic-Quant-GGUF install, users can directly use Qwen3.5-PRISM-Dynamic-Quant-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Qwen3.5-PRISM-Dynamic-Quant-GGUF install url in huggingface.co: