internlm / internlm2_5-7b-chat-4bit

huggingface.co
Total runs: 248
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: July 02 2024
text-generation

Introduction of internlm2_5-7b-chat-4bit

Model Details of internlm2_5-7b-chat-4bit

INT4 Weight-only Quantization and Deployment (W4A16)

LMDeploy adopts AWQ algorithm for 4bit weight-only quantization. By developed the high-performance cuda kernel, the 4bit quantized model inference achieves up to 2.4x faster than FP16.

LMDeploy supports the following NVIDIA GPU for W4A16 inference:

  • Turing(sm75): 20 series, T4

  • Ampere(sm80,sm86): 30 series, A10, A16, A30, A100

  • Ada Lovelace(sm90): 40 series

Before proceeding with the quantization and inference, please ensure that lmdeploy is installed.

pip install lmdeploy[all]

This article comprises the following sections:

Inference

Trying the following codes, you can perform the batched offline inference with the quantized model:

from lmdeploy import pipeline, TurbomindEngineConfig
engine_config = TurbomindEngineConfig(model_format='awq')
pipe = pipeline("internlm/internlm2_5-7b-chat-4bit", backend_config=engine_config)
response = pipe(["Hi, pls intro yourself", "Shanghai is"])
print(response)

For more information about the pipeline parameters, please refer to here .

Evaluation

Please overview this guide about model evaluation with LMDeploy.

Service

LMDeploy's api_server enables models to be easily packed into services with a single command. The provided RESTful APIs are compatible with OpenAI's interfaces. Below are an example of service startup:

lmdeploy serve api_server internlm/internlm2_5-7b-chat-4bit --backend turbomind --model-format awq

The default port of api_server is 23333 . After the server is launched, you can communicate with server on terminal through api_client :

lmdeploy serve api_client http://0.0.0.0:23333

You can overview and try out api_server APIs online by swagger UI at http://0.0.0.0:23333 , or you can also read the API specification from here .

Runs of internlm internlm2_5-7b-chat-4bit on huggingface.co

248
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About internlm2_5-7b-chat-4bit huggingface.co Model

More internlm2_5-7b-chat-4bit license Visit here:

https://choosealicense.com/licenses/apache-2.0

internlm2_5-7b-chat-4bit huggingface.co

internlm2_5-7b-chat-4bit huggingface.co is an AI model on huggingface.co that provides internlm2_5-7b-chat-4bit's model effect (), which can be used instantly with this internlm internlm2_5-7b-chat-4bit model. huggingface.co supports a free trial of the internlm2_5-7b-chat-4bit model, and also provides paid use of the internlm2_5-7b-chat-4bit. Support call internlm2_5-7b-chat-4bit model through api, including Node.js, Python, http.

internlm2_5-7b-chat-4bit huggingface.co Url

https://huggingface.co/internlm/internlm2_5-7b-chat-4bit

internlm internlm2_5-7b-chat-4bit online free

internlm2_5-7b-chat-4bit huggingface.co is an online trial and call api platform, which integrates internlm2_5-7b-chat-4bit's modeling effects, including api services, and provides a free online trial of internlm2_5-7b-chat-4bit, you can try internlm2_5-7b-chat-4bit online for free by clicking the link below.

internlm internlm2_5-7b-chat-4bit online free url in huggingface.co:

https://huggingface.co/internlm/internlm2_5-7b-chat-4bit

internlm2_5-7b-chat-4bit install

internlm2_5-7b-chat-4bit is an open source model from GitHub that offers a free installation service, and any user can find internlm2_5-7b-chat-4bit on GitHub to install. At the same time, huggingface.co provides the effect of internlm2_5-7b-chat-4bit install, users can directly use internlm2_5-7b-chat-4bit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

internlm2_5-7b-chat-4bit install url in huggingface.co:

https://huggingface.co/internlm/internlm2_5-7b-chat-4bit

Url of internlm2_5-7b-chat-4bit

internlm2_5-7b-chat-4bit huggingface.co Url

Provider of internlm2_5-7b-chat-4bit huggingface.co

internlm
ORGANIZATIONS

Other API from internlm

huggingface.co

Total runs: 25.1K
Run Growth: 2.9K
Growth Rate: 11.51%
Updated:March 13 2025
huggingface.co

Total runs: 21.4K
Run Growth: -29.4K
Growth Rate: -137.33%
Updated:March 29 2026
huggingface.co

Total runs: 2.8K
Run Growth: 2.3K
Growth Rate: 81.18%
Updated:July 03 2024
huggingface.co

Total runs: 327
Run Growth: 35
Growth Rate: 10.70%
Updated:April 16 2026
huggingface.co

Total runs: 123
Run Growth: 91
Growth Rate: 73.98%
Updated:October 23 2025