Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
ParoQuant is the state-of-the-art INT4 quantization for LLMs. It closes the accuracy gap with FP16 while running at near-AWQ speed. Supports NVIDIA GPUs (vLLM, Transformers) and Apple Silicon (MLX).
# Interactive chat
docker run --pull=always --rm -it --gpus all --ipc=host \
ghcr.io/z-lab/paroquant:chat --model z-lab/Llama-3.1-8B-Instruct-PARO
# API server (port 8000)
docker run --pull=always --rm -it --gpus all --ipc=host -p 8000:8000 \
ghcr.io/z-lab/paroquant:serve --model z-lab/Llama-3.1-8B-Instruct-PARO
Citation
@inproceedings{liang2026paroquant,
title = {{ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference}},
author = {Liang, Yesheng and Chen, Haisheng and Zhang, Zihan and Han, Song and Liu, Zhijian},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}
Runs of z-lab Llama-3.1-8B-Instruct-PARO on huggingface.co
1.4K
Total runs
0
24-hour runs
47
3-day runs
401
7-day runs
1.2K
30-day runs
More Information About Llama-3.1-8B-Instruct-PARO huggingface.co Model
More Llama-3.1-8B-Instruct-PARO license Visit here:
Llama-3.1-8B-Instruct-PARO huggingface.co is an AI model on huggingface.co that provides Llama-3.1-8B-Instruct-PARO's model effect (), which can be used instantly with this z-lab Llama-3.1-8B-Instruct-PARO model. huggingface.co supports a free trial of the Llama-3.1-8B-Instruct-PARO model, and also provides paid use of the Llama-3.1-8B-Instruct-PARO. Support call Llama-3.1-8B-Instruct-PARO model through api, including Node.js, Python, http.
Llama-3.1-8B-Instruct-PARO huggingface.co is an online trial and call api platform, which integrates Llama-3.1-8B-Instruct-PARO's modeling effects, including api services, and provides a free online trial of Llama-3.1-8B-Instruct-PARO, you can try Llama-3.1-8B-Instruct-PARO online for free by clicking the link below.
z-lab Llama-3.1-8B-Instruct-PARO online free url in huggingface.co:
Llama-3.1-8B-Instruct-PARO is an open source model from GitHub that offers a free installation service, and any user can find Llama-3.1-8B-Instruct-PARO on GitHub to install. At the same time, huggingface.co provides the effect of Llama-3.1-8B-Instruct-PARO install, users can directly use Llama-3.1-8B-Instruct-PARO installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Llama-3.1-8B-Instruct-PARO install url in huggingface.co: