We have optimized the storage method of the KV Cache, reducing the generation of memory fragmentation. Based on the optimized code, the model can process a context length of 32K under approximately
20G
of memory (FP/BF16 format).
The ChatGLM
2
-6B-32K further strengthens the ability to understand long texts based on the
ChatGLM2-6B
, and can better handle up to 32K context length. Specifically, we have updated the position encoding based on the method of
Positional Interpolation
, and trained with a 32K context length during the dialogue alignment. In practical use, if the context length you are dealing with is generally within 8K, we recommend using
ChatGLM2-6B
; if you need to handle a context length exceeding 8K, we recommend using ChatGLM2-6B-32K.
ChatGLM2-6B-32K is the second-generation version of the open-source bilingual (Chinese-English) chat model
ChatGLM-6B
. It retains the smooth conversation flow and low deployment threshold of the first-generation model, while introducing the following new features:
Stronger Performance
: Based on the development experience of the first-generation ChatGLM model, we have fully upgraded the base model of ChatGLM2-6B-32K. ChatGLM2-6B-32K uses the hybrid objective function of
GLM
, and has undergone pre-training with 1.4T bilingual tokens and human preference alignment training.
Longer Context
: Based on
FlashAttention
technique, we have extended the context length of the base model from 2K in ChatGLM-6B to 32K, and trained with a context length of 32K during the dialogue alignment, allowing for more rounds of dialogue.
More Efficient Inference
: Based on
Multi-Query Attention
technique, ChatGLM2-6B-32K has more efficient inference speed and lower GPU memory usage: under the official implementation, the inference speed has increased by 42% compared to the first generation; under INT4 quantization, the dialogue length supported by 6G GPU memory has increased from 1K to 8K.
More Open License
: ChatGLM2-6B-32K weights are
completely open
for academic research, and
free commercial use
is also allowed after completing the
questionnaire
.
If you find our work helpful, please consider citing the following paper.
@misc{glm2024chatglm,
title={ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools},
author={Team GLM and Aohan Zeng and Bin Xu and Bowen Wang and Chenhui Zhang and Da Yin and Diego Rojas and Guanyu Feng and Hanlin Zhao and Hanyu Lai and Hao Yu and Hongning Wang and Jiadai Sun and Jiajie Zhang and Jiale Cheng and Jiayi Gui and Jie Tang and Jing Zhang and Juanzi Li and Lei Zhao and Lindong Wu and Lucen Zhong and Mingdao Liu and Minlie Huang and Peng Zhang and Qinkai Zheng and Rui Lu and Shuaiqi Duan and Shudan Zhang and Shulin Cao and Shuxun Yang and Weng Lam Tam and Wenyi Zhao and Xiao Liu and Xiao Xia and Xiaohan Zhang and Xiaotao Gu and Xin Lv and Xinghan Liu and Xinyi Liu and Xinyue Yang and Xixuan Song and Xunkai Zhang and Yifan An and Yifan Xu and Yilin Niu and Yuantao Yang and Yueyan Li and Yushi Bai and Yuxiao Dong and Zehan Qi and Zhaoyu Wang and Zhen Yang and Zhengxiao Du and Zhenyu Hou and Zihan Wang},
year={2024},
eprint={2406.12793},
archivePrefix={arXiv},
primaryClass={id='cs.CL' full_name='Computation and Language' is_active=True alt_name='cmp-lg' in_archive='cs' is_general=False description='Covers natural language processing. Roughly includes material in ACM Subject Class I.2.7. Note that work on artificial languages (programming languages, logics, formal systems) that does not explicitly address natural-language issues broadly construed (natural-language processing, computational linguistics, speech, text retrieval, etc.) is not appropriate for this area.'}
}
Runs of zai-org chatglm2-6b-32k on huggingface.co
264
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
2
30-day runs
More Information About chatglm2-6b-32k huggingface.co Model
chatglm2-6b-32k huggingface.co
chatglm2-6b-32k huggingface.co is an AI model on huggingface.co that provides chatglm2-6b-32k's model effect (), which can be used instantly with this zai-org chatglm2-6b-32k model. huggingface.co supports a free trial of the chatglm2-6b-32k model, and also provides paid use of the chatglm2-6b-32k. Support call chatglm2-6b-32k model through api, including Node.js, Python, http.
chatglm2-6b-32k huggingface.co is an online trial and call api platform, which integrates chatglm2-6b-32k's modeling effects, including api services, and provides a free online trial of chatglm2-6b-32k, you can try chatglm2-6b-32k online for free by clicking the link below.
zai-org chatglm2-6b-32k online free url in huggingface.co:
chatglm2-6b-32k is an open source model from GitHub that offers a free installation service, and any user can find chatglm2-6b-32k on GitHub to install. At the same time, huggingface.co provides the effect of chatglm2-6b-32k install, users can directly use chatglm2-6b-32k installed effect in huggingface.co for debugging and trial. It also supports api for free installation.