A bidirectional language model based on the Encoder structure, focusing on solving various NLU tasks.
We follow
Megatron-LM
, using 32 A100s and spending 14 days training a billion-level BERT on WuDao Corpora (180 GB version). Given Chinese grammar and the difficulty of large-scale training, we use four pre-training procedures to improve BERT: 1) Whole Word Masking (WWM), 2) Knowledge-based Dynamic Masking (KDM), 3) Sentence Order Prediction (SOP), 4) Pre-layer Normalization (Pre-LN).
1.On November 10, 2021, Erlangshen-MegatronBert-1.3B topped the FewCLUE benchmark. Among them, our Erlangshen outperformed human performance in CHIDF (idiom fill-in-the-blank) and TNEWS (news classification) subtasks. In addition, our Erlangshen ranked the top in CHIDF (idiom fill-in-the-blank), CSLDCP (subject literature classification), and OCNLI (natural language inference) tasks.
2.On January 24, 2022, Erlangshen-MegatronBert-1.3B topped the ZeroCLUE benchmark. For each of these tasks, we rank the top ones in CSLDCP (Subject Literature Classification), TNEWS (News Classification), IFLYTEK (Application Description Classification), CSL (Abstract Keyword Recognition), and CLUEWSC (Referential Disambiguation) tasks.
3.Erlangshen-MegatronBert-1.3B topped the CLUE benchmark semantic matching task on July 10, 2022.
下游效果 Performance
模型
afqmc
tnews
iflytek
ocnli
cmnli
wsc
csl
roberta-wwm-ext-large
0.7514
0.5872
0.6152
0.777
0.814
0.8914
0.86
Erlangshen-MegatronBert-1.3B
0.7608
0.5996
0.6234
0.7917
0.81
0.9243
0.872
使用 Usage
from transformers import MegatronBertConfig, MegatronBertModel
from transformers import BertTokenizer
tokenizer = BertTokenizer.from_pretrained("IDEA-CCNL/Erlangshen-MegatronBert-1.3B")
config = MegatronBertConfig.from_pretrained("IDEA-CCNL/Erlangshen-MegatronBert-1.3B")
model = MegatronBertModel.from_pretrained("IDEA-CCNL/Erlangshen-MegatronBert-1.3B")
If you are using the resource for your work, please cite the our
paper
:
@article{fengshenbang,
author = {Jiaxing Zhang and Ruyi Gan and Junjie Wang and Yuxiang Zhang and Lin Zhang and Ping Yang and Xinyu Gao and Ziwei Wu and Xiaoqun Dong and Junqing He and Jianheng Zhuo and Qi Yang and Yongfeng Huang and Xiayu Li and Yanghan Wu and Junyu Lu and Xinyu Zhu and Weifeng Chen and Ting Han and Kunhao Pan and Rui Wang and Hao Wang and Xiaojun Wu and Zhongshen Zeng and Chongpei Chen},
title = {Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence},
journal = {CoRR},
volume = {abs/2209.02970},
year = {2022}
}
Erlangshen-MegatronBert-1.3B huggingface.co is an AI model on huggingface.co that provides Erlangshen-MegatronBert-1.3B's model effect (), which can be used instantly with this IDEA-CCNL Erlangshen-MegatronBert-1.3B model. huggingface.co supports a free trial of the Erlangshen-MegatronBert-1.3B model, and also provides paid use of the Erlangshen-MegatronBert-1.3B. Support call Erlangshen-MegatronBert-1.3B model through api, including Node.js, Python, http.
Erlangshen-MegatronBert-1.3B huggingface.co is an online trial and call api platform, which integrates Erlangshen-MegatronBert-1.3B's modeling effects, including api services, and provides a free online trial of Erlangshen-MegatronBert-1.3B, you can try Erlangshen-MegatronBert-1.3B online for free by clicking the link below.
IDEA-CCNL Erlangshen-MegatronBert-1.3B online free url in huggingface.co:
Erlangshen-MegatronBert-1.3B is an open source model from GitHub that offers a free installation service, and any user can find Erlangshen-MegatronBert-1.3B on GitHub to install. At the same time, huggingface.co provides the effect of Erlangshen-MegatronBert-1.3B install, users can directly use Erlangshen-MegatronBert-1.3B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Erlangshen-MegatronBert-1.3B install url in huggingface.co: