MiniCPM is an End-Size LLM developed by ModelBest Inc. and TsinghuaNLP, with only 2.4B parameters excluding embeddings.
MiniCPM-2B-128k is a long context extension trial of
MiniCPM-2B
.
To our best knowledge, MiniCPM-2B-128k is the first long context(>=128k) SLM smaller than 3B。
In comparison with the previous released
MiniCPM-2B
, the improvements include:
Supports 128k context, achieving the best score under 7B on the comprehensive long-text evaluation InfiniteBench, but performance drops within 4k context
To facilitate community developers, the model has updated the {} directive template to chatml format (user\n{}\nassistant\n) during alignment, which also aids users in deploying and using the vllm openai compatible server mode.
Due to the parallel mechanism requirement, removed tie_embedding and expanded the vocabulary to 127660.
Notice: We discovered that the quality of Huggingface generation is slightly lower and significantly slower than vLLM, thus benchmarking using vLLM is recommended.
注意:我们发现使用Huggingface生成质量略差于vLLM,因此推荐使用vLLM进行测试。
Limitations 局限性
Due to limitations in model size, the model may experience hallucinatory issues. As DPO model tend to generate longer response, hallucinations are more likely to occur. We will also continue to iterate and improve the MiniCPM model.
To ensure the universality of the model for academic research purposes, we did not conduct any identity training on the model. Meanwhile, as we use ShareGPT open-source corpus as part of the training data, the model may output identity information similar to the GPT series models.
Due to the limitation of model size, the output of the model is greatly influenced by prompt words, which may result in inconsistent results from multiple attempts.
Due to limited model capacity, the model's knowledge memory is not accurate. In the future, we will combine the RAG method to enhance the model's knowledge memory ability.
Run the following code after install transformers>=4.36.0 and accelerate
Warning: It is necessary to specify the data type of the model clearly in 'from_pretrained', otherwise large calculation errors will be caused
安装transformers>=4.36.0以及accelerate后,运行以下代码
注意:需要在from_pretrained中明确指明模型的数据类型,否则会引起较大计算误差
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
torch.manual_seed(0)
path = 'openbmb/MiniCPM-2B-128k'
tokenizer = AutoTokenizer.from_pretrained(path)
model = AutoModelForCausalLM.from_pretrained(path, torch_dtype=torch.bfloat16, device_map='cuda', trust_remote_code=True)
responds, history = model.chat(tokenizer, "山东省最高的山是哪座山, 它比黄山高还是矮?差距多少?", temperature=0.8, top_p=0.8)
print(responds)
Runs of openbmb MiniCPM-2B-128k on huggingface.co
1.1K
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
-109
30-day runs
More Information About MiniCPM-2B-128k huggingface.co Model
MiniCPM-2B-128k huggingface.co
MiniCPM-2B-128k huggingface.co is an AI model on huggingface.co that provides MiniCPM-2B-128k's model effect (), which can be used instantly with this openbmb MiniCPM-2B-128k model. huggingface.co supports a free trial of the MiniCPM-2B-128k model, and also provides paid use of the MiniCPM-2B-128k. Support call MiniCPM-2B-128k model through api, including Node.js, Python, http.
MiniCPM-2B-128k huggingface.co is an online trial and call api platform, which integrates MiniCPM-2B-128k's modeling effects, including api services, and provides a free online trial of MiniCPM-2B-128k, you can try MiniCPM-2B-128k online for free by clicking the link below.
openbmb MiniCPM-2B-128k online free url in huggingface.co:
MiniCPM-2B-128k is an open source model from GitHub that offers a free installation service, and any user can find MiniCPM-2B-128k on GitHub to install. At the same time, huggingface.co provides the effect of MiniCPM-2B-128k install, users can directly use MiniCPM-2B-128k installed effect in huggingface.co for debugging and trial. It also supports api for free installation.