internlm / internlm2-chat-7b-4bits

huggingface.co
Total runs: 109
24-hour runs: 0
7-day runs: -6
30-day runs: 22
Model's Last Updated: April 24 2024
text-generation

Introduction of internlm2-chat-7b-4bits

Model Details of internlm2-chat-7b-4bits

INT4 Weight-only Quantization and Deployment (W4A16)

LMDeploy adopts AWQ algorithm for 4bit weight-only quantization. By developed the high-performance cuda kernel, the 4bit quantized model inference achieves up to 2.4x faster than FP16.

LMDeploy supports the following NVIDIA GPU for W4A16 inference:

  • Turing(sm75): 20 series, T4

  • Ampere(sm80,sm86): 30 series, A10, A16, A30, A100

  • Ada Lovelace(sm90): 40 series

Before proceeding with the quantization and inference, please ensure that lmdeploy is installed.

pip install lmdeploy[all]

This article comprises the following sections:

Inference

Trying the following codes, you can perform the batched offline inference with the quantized model:

from lmdeploy import pipeline, TurbomindEngineConfig
engine_config = TurbomindEngineConfig(model_format='awq')
pipe = pipeline("internlm/internlm2-chat-7b-4bits", backend_config=engine_config)
response = pipe(["Hi, pls intro yourself", "Shanghai is"])
print(response)

For more information about the pipeline parameters, please refer to here .

Evaluation

Please overview this guide about model evaluation with LMDeploy.

Service

LMDeploy's api_server enables models to be easily packed into services with a single command. The provided RESTful APIs are compatible with OpenAI's interfaces. Below are an example of service startup:

lmdeploy serve api_server internlm/internlm2-chat-7b-4bits --backend turbomind --model-format awq

The default port of api_server is 23333 . After the server is launched, you can communicate with server on terminal through api_client :

lmdeploy serve api_client http://0.0.0.0:23333

You can overview and try out api_server APIs online by swagger UI at http://0.0.0.0:23333 , or you can also read the API specification from here .

Runs of internlm internlm2-chat-7b-4bits on huggingface.co

109
Total runs
0
24-hour runs
-6
3-day runs
-6
7-day runs
22
30-day runs

More Information About internlm2-chat-7b-4bits huggingface.co Model

More internlm2-chat-7b-4bits license Visit here:

https://choosealicense.com/licenses/apache-2.0

internlm2-chat-7b-4bits huggingface.co

internlm2-chat-7b-4bits huggingface.co is an AI model on huggingface.co that provides internlm2-chat-7b-4bits's model effect (), which can be used instantly with this internlm internlm2-chat-7b-4bits model. huggingface.co supports a free trial of the internlm2-chat-7b-4bits model, and also provides paid use of the internlm2-chat-7b-4bits. Support call internlm2-chat-7b-4bits model through api, including Node.js, Python, http.

internlm2-chat-7b-4bits huggingface.co Url

https://huggingface.co/internlm/internlm2-chat-7b-4bits

internlm internlm2-chat-7b-4bits online free

internlm2-chat-7b-4bits huggingface.co is an online trial and call api platform, which integrates internlm2-chat-7b-4bits's modeling effects, including api services, and provides a free online trial of internlm2-chat-7b-4bits, you can try internlm2-chat-7b-4bits online for free by clicking the link below.

internlm internlm2-chat-7b-4bits online free url in huggingface.co:

https://huggingface.co/internlm/internlm2-chat-7b-4bits

internlm2-chat-7b-4bits install

internlm2-chat-7b-4bits is an open source model from GitHub that offers a free installation service, and any user can find internlm2-chat-7b-4bits on GitHub to install. At the same time, huggingface.co provides the effect of internlm2-chat-7b-4bits install, users can directly use internlm2-chat-7b-4bits installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

internlm2-chat-7b-4bits install url in huggingface.co:

https://huggingface.co/internlm/internlm2-chat-7b-4bits

Url of internlm2-chat-7b-4bits

internlm2-chat-7b-4bits huggingface.co Url

Provider of internlm2-chat-7b-4bits huggingface.co

internlm
ORGANIZATIONS

Other API from internlm

huggingface.co

Total runs: 21.4K
Run Growth: -29.4K
Growth Rate: -137.33%
Updated:March 29 2026
huggingface.co

Total runs: 2.8K
Run Growth: 2.3K
Growth Rate: 82.73%
Updated:July 03 2024
huggingface.co

Total runs: 327
Run Growth: 35
Growth Rate: 10.70%
Updated:April 16 2026
huggingface.co

Total runs: 123
Run Growth: 91
Growth Rate: 73.98%
Updated:October 23 2025