We have optimized the storage method of the KV Cache, reducing the generation of memory fragmentation. Based on the optimized code, the model can process a context length of 32K under approximately
11G
of memory.
The ChatGLM
2
-6B-32K further strengthens the ability to understand long texts based on the
ChatGLM2-6B
, and can better handle up to 32K context length. Specifically, we have updated the position encoding based on the method of
Positional Interpolation
, and trained with a 32K context length during the dialogue alignment. In practical use, if the context length you are dealing with is generally within 8K, we recommend using
ChatGLM2-6B
; if you need to handle a context length exceeding 8K, we recommend using ChatGLM2-6B-32K.
ChatGLM2-6B-32K is the second-generation version of the open-source bilingual (Chinese-English) chat model
ChatGLM-6B
. It retains the smooth conversation flow and low deployment threshold of the first-generation model, while introducing the following new features:
Stronger Performance
: Based on the development experience of the first-generation ChatGLM model, we have fully upgraded the base model of ChatGLM2-6B-32K. ChatGLM2-6B-32K uses the hybrid objective function of
GLM
, and has undergone pre-training with 1.4T bilingual tokens and human preference alignment training.
Longer Context
: Based on
FlashAttention
technique, we have extended the context length of the base model from 2K in ChatGLM-6B to 32K, and trained with a context length of 32K during the dialogue alignment, allowing for more rounds of dialogue.
More Efficient Inference
: Based on
Multi-Query Attention
technique, ChatGLM2-6B-32K has more efficient inference speed and lower GPU memory usage: under the official implementation, the inference speed has increased by 42% compared to the first generation; under INT4 quantization, the dialogue length supported by 6G GPU memory has increased from 1K to 8K.
More Open License
: ChatGLM2-6B-32K weights are
completely open
for academic research, and
free commercial use
is also allowed after completing the
questionnaire
.
@article{zeng2022glm,
title={Glm-130b: An open bilingual pre-trained model},
author={Zeng, Aohan and Liu, Xiao and Du, Zhengxiao and Wang, Zihan and Lai, Hanyu and Ding, Ming and Yang, Zhuoyi and Xu, Yifan and Zheng, Wendi and Xia, Xiao and others},
journal={arXiv preprint arXiv:2210.02414},
year={2022}
}
@inproceedings{du2022glm,
title={GLM: General Language Model Pretraining with Autoregressive Blank Infilling},
author={Du, Zhengxiao and Qian, Yujie and Liu, Xiao and Ding, Ming and Qiu, Jiezhong and Yang, Zhilin and Tang, Jie},
booktitle={Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
pages={320--335},
year={2022}
}
Runs of THUDM chatglm2-6b-32k-int4 on huggingface.co
15
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About chatglm2-6b-32k-int4 huggingface.co Model
chatglm2-6b-32k-int4 huggingface.co
chatglm2-6b-32k-int4 huggingface.co is an AI model on huggingface.co that provides chatglm2-6b-32k-int4's model effect (), which can be used instantly with this THUDM chatglm2-6b-32k-int4 model. huggingface.co supports a free trial of the chatglm2-6b-32k-int4 model, and also provides paid use of the chatglm2-6b-32k-int4. Support call chatglm2-6b-32k-int4 model through api, including Node.js, Python, http.
chatglm2-6b-32k-int4 huggingface.co is an online trial and call api platform, which integrates chatglm2-6b-32k-int4's modeling effects, including api services, and provides a free online trial of chatglm2-6b-32k-int4, you can try chatglm2-6b-32k-int4 online for free by clicking the link below.
THUDM chatglm2-6b-32k-int4 online free url in huggingface.co:
chatglm2-6b-32k-int4 is an open source model from GitHub that offers a free installation service, and any user can find chatglm2-6b-32k-int4 on GitHub to install. At the same time, huggingface.co provides the effect of chatglm2-6b-32k-int4 install, users can directly use chatglm2-6b-32k-int4 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
chatglm2-6b-32k-int4 install url in huggingface.co: