Introduction of roberta-tiny-word-chinese-cluecorpussmall
Model Details of roberta-tiny-word-chinese-cluecorpussmall
Chinese word-based RoBERTa Miniatures
Model description
This is the set of 5 Chinese word-based RoBERTa models pre-trained by
UER-py
, which is introduced in
this paper
. Besides, the models could also be pre-trained by
TencentPretrain
introduced in
this paper
, which inherits UER-py to support models with parameters above one billion, and extends it to a multimodal pre-training framework.
Most Chinese pre-trained weights are based on Chinese character. Compared with character-based models, word-based models are faster (because of shorter sequence length) and have better performance according to our experimental results. To this end, we released the 5 Chinese word-based RoBERTa models of different sizes. In order to facilitate users in reproducing the results, we used a publicly available corpus and word segmentation tool, and provided all training details.
You can download the 5 Chinese RoBERTa miniatures either from the
UER-py Modelzoo page
, or via HuggingFace from the links below:
Here is how to use this model to get the features of a given text in PyTorch:
from transformers import AlbertTokenizer, BertModel
tokenizer = AlbertTokenizer.from_pretrained('uer/roberta-medium-word-chinese-cluecorpussmall')
model = BertModel.from_pretrained("uer/roberta-medium-word-chinese-cluecorpussmall")
text = "用你喜欢的任何文本替换我。"
encoded_input = tokenizer(text, return_tensors='pt')
output = model(**encoded_input)
and in TensorFlow:
from transformers import AlbertTokenizer, TFBertModel
tokenizer = AlbertTokenizer.from_pretrained('uer/roberta-medium-word-chinese-cluecorpussmall')
model = TFBertModel.from_pretrained("uer/roberta-medium-word-chinese-cluecorpussmall")
text = "用你喜欢的任何文本替换我。"
encoded_input = tokenizer(text, return_tensors='tf')
output = model(encoded_input)
Since BertTokenizer does not support sentencepiece, AlbertTokenizer is used here.
Training data
CLUECorpusSmall
is used as training data. Google's
sentencepiece
is used for word segmentation. The sentencepiece model is trained on CLUECorpusSmall corpus:
Models are pre-trained by
UER-py
on
Tencent Cloud
. We pre-train 1,000,000 steps with a sequence length of 128 and then pre-train 250,000 additional steps with a sequence length of 512. We use the same hyper-parameters on different model sizes.
@article{devlin2018bert,
title={BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding},
author={Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},
journal={arXiv preprint arXiv:1810.04805},
year={2018}
}
@article{turc2019,
title={Well-Read Students Learn Better: On the Importance of Pre-training Compact Models},
author={Turc, Iulia and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},
journal={arXiv preprint arXiv:1908.08962v2 },
year={2019}
}
@article{zhao2019uer,
title={UER: An Open-Source Toolkit for Pre-training Models},
author={Zhao, Zhe and Chen, Hui and Zhang, Jinbin and Zhao, Xin and Liu, Tao and Lu, Wei and Chen, Xi and Deng, Haotang and Ju, Qi and Du, Xiaoyong},
journal={EMNLP-IJCNLP 2019},
pages={241},
year={2019}
}
@article{zhao2023tencentpretrain,
title={TencentPretrain: A Scalable and Flexible Toolkit for Pre-training Models of Different Modalities},
author={Zhao, Zhe and Li, Yudong and Hou, Cheng and Zhao, Jing and others},
journal={ACL 2023},
pages={217},
year={2023}
Runs of uer roberta-tiny-word-chinese-cluecorpussmall on huggingface.co
87
Total runs
1
24-hour runs
1
3-day runs
6
7-day runs
40
30-day runs
More Information About roberta-tiny-word-chinese-cluecorpussmall huggingface.co Model
roberta-tiny-word-chinese-cluecorpussmall huggingface.co is an AI model on huggingface.co that provides roberta-tiny-word-chinese-cluecorpussmall's model effect (), which can be used instantly with this uer roberta-tiny-word-chinese-cluecorpussmall model. huggingface.co supports a free trial of the roberta-tiny-word-chinese-cluecorpussmall model, and also provides paid use of the roberta-tiny-word-chinese-cluecorpussmall. Support call roberta-tiny-word-chinese-cluecorpussmall model through api, including Node.js, Python, http.
roberta-tiny-word-chinese-cluecorpussmall huggingface.co is an online trial and call api platform, which integrates roberta-tiny-word-chinese-cluecorpussmall's modeling effects, including api services, and provides a free online trial of roberta-tiny-word-chinese-cluecorpussmall, you can try roberta-tiny-word-chinese-cluecorpussmall online for free by clicking the link below.
uer roberta-tiny-word-chinese-cluecorpussmall online free url in huggingface.co:
roberta-tiny-word-chinese-cluecorpussmall is an open source model from GitHub that offers a free installation service, and any user can find roberta-tiny-word-chinese-cluecorpussmall on GitHub to install. At the same time, huggingface.co provides the effect of roberta-tiny-word-chinese-cluecorpussmall install, users can directly use roberta-tiny-word-chinese-cluecorpussmall installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
roberta-tiny-word-chinese-cluecorpussmall install url in huggingface.co: