INT4 Weight-only Quantization and Deployment (W4A16)
LMDeploy adopts
AWQ
algorithm for 4bit weight-only quantization. By developed the high-performance cuda kernel, the 4bit quantized model inference achieves up to 2.4x faster than FP16.
LMDeploy supports the following NVIDIA GPU for W4A16 inference:
Turing(sm75): 20 series, T4
Ampere(sm80,sm86): 30 series, A10, A16, A30, A100
Ada Lovelace(sm90): 40 series
Before proceeding with the quantization and inference, please ensure that lmdeploy is installed.
For more information about the pipeline parameters, please refer to
here
.
Evaluation
Please overview
this guide
about model evaluation with LMDeploy.
Service
LMDeploy's
api_server
enables models to be easily packed into services with a single command. The provided RESTful APIs are compatible with OpenAI's interfaces. Below are an example of service startup:
internlm2-chat-7b-4bits huggingface.co is an AI model on huggingface.co that provides internlm2-chat-7b-4bits's model effect (), which can be used instantly with this internlm internlm2-chat-7b-4bits model. huggingface.co supports a free trial of the internlm2-chat-7b-4bits model, and also provides paid use of the internlm2-chat-7b-4bits. Support call internlm2-chat-7b-4bits model through api, including Node.js, Python, http.
internlm2-chat-7b-4bits huggingface.co is an online trial and call api platform, which integrates internlm2-chat-7b-4bits's modeling effects, including api services, and provides a free online trial of internlm2-chat-7b-4bits, you can try internlm2-chat-7b-4bits online for free by clicking the link below.
internlm internlm2-chat-7b-4bits online free url in huggingface.co:
internlm2-chat-7b-4bits is an open source model from GitHub that offers a free installation service, and any user can find internlm2-chat-7b-4bits on GitHub to install. At the same time, huggingface.co provides the effect of internlm2-chat-7b-4bits install, users can directly use internlm2-chat-7b-4bits installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
internlm2-chat-7b-4bits install url in huggingface.co: