z-lab / Qwen3-8B-PARO

huggingface.co
Total runs: 339
24-hour runs: 0
7-day runs: -20
30-day runs: -20
Model's Last Updated: May 08 2026
text-generation

Introduction of Qwen3-8B-PARO

Model Details of Qwen3-8B-PARO

z-lab/Qwen3-8B-PARO

Pairwise Rotation Quantization for Efficient Reasoning LLM Inference

Paper Blog Models PyPI

ParoQuant is the state-of-the-art INT4 quantization for LLMs. It closes the accuracy gap with FP16 while running at near-AWQ speed. Supports NVIDIA GPUs (vLLM, Transformers) and Apple Silicon (MLX).

z-lab/Qwen3-8B-PARO is a 4-bit Qwen/Qwen3-8B quantized with ParoQuant . Check out other ParoQuant models from the Hugging Face collection . Swap the model name in the commands below to try any of them.

Quick Start
Installation
# NVIDIA GPU
pip install "paroquant[vllm]"

# Apple Silicon
pip install "paroquant[mlx]"
Interactive Chat
python -m paroquant.cli.chat --model z-lab/Qwen3-8B-PARO
OpenAI-Compatible API Server
python -m paroquant.cli.serve --model z-lab/Qwen3-8B-PARO --port 8000
Agent with Tool Calling

Start the API server first, then install the agent dependencies and run:

pip install "paroquant[agent]"
python -m paroquant.cli.agent --model z-lab/Qwen3-8B-PARO

Tool use (web fetch, filesystem, time) requires Node.js .

Docker (NVIDIA GPU)
# Interactive chat
docker run --pull=always --rm -it --gpus all --ipc=host \
  ghcr.io/z-lab/paroquant:chat --model z-lab/Qwen3-8B-PARO

# API server (port 8000)
docker run --pull=always --rm -it --gpus all --ipc=host -p 8000:8000 \
  ghcr.io/z-lab/paroquant:serve --model z-lab/Qwen3-8B-PARO
Citation
@inproceedings{liang2026paroquant,
  title     = {{ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference}},
  author    = {Liang, Yesheng and Chen, Haisheng and Zhang, Zihan and Han, Song and Liu, Zhijian},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2026}
}

Runs of z-lab Qwen3-8B-PARO on huggingface.co

339
Total runs
0
24-hour runs
12
3-day runs
-20
7-day runs
-20
30-day runs

More Information About Qwen3-8B-PARO huggingface.co Model

More Qwen3-8B-PARO license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen3-8B-PARO huggingface.co

Qwen3-8B-PARO huggingface.co is an AI model on huggingface.co that provides Qwen3-8B-PARO's model effect (), which can be used instantly with this z-lab Qwen3-8B-PARO model. huggingface.co supports a free trial of the Qwen3-8B-PARO model, and also provides paid use of the Qwen3-8B-PARO. Support call Qwen3-8B-PARO model through api, including Node.js, Python, http.

Qwen3-8B-PARO huggingface.co Url

https://huggingface.co/z-lab/Qwen3-8B-PARO

z-lab Qwen3-8B-PARO online free

Qwen3-8B-PARO huggingface.co is an online trial and call api platform, which integrates Qwen3-8B-PARO's modeling effects, including api services, and provides a free online trial of Qwen3-8B-PARO, you can try Qwen3-8B-PARO online for free by clicking the link below.

z-lab Qwen3-8B-PARO online free url in huggingface.co:

https://huggingface.co/z-lab/Qwen3-8B-PARO

Qwen3-8B-PARO install

Qwen3-8B-PARO is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-8B-PARO on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-8B-PARO install, users can directly use Qwen3-8B-PARO installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3-8B-PARO install url in huggingface.co:

https://huggingface.co/z-lab/Qwen3-8B-PARO

Url of Qwen3-8B-PARO

Qwen3-8B-PARO huggingface.co Url

Provider of Qwen3-8B-PARO huggingface.co

z-lab
ORGANIZATIONS

Other API from z-lab

huggingface.co

Total runs: 1.9K
Run Growth: -15.0K
Growth Rate: -800.00%
Updated:May 08 2026
huggingface.co

Total runs: 311
Run Growth: -42
Growth Rate: -13.50%
Updated:May 08 2026
huggingface.co

Total runs: 242
Run Growth: -111
Growth Rate: -45.87%
Updated:May 08 2026
huggingface.co

Total runs: 9
Run Growth: 0
Growth Rate: 0.00%
Updated:June 23 2025