BERTweet: A pre-trained language model for English Tweets
BERTweet is the first public large-scale language model pre-trained for English Tweets. BERTweet is trained based on the
RoBERTa
pre-training procedure. The corpus used to pre-train BERTweet consists of 850M English Tweets (16B word tokens ~ 80GB), containing 845M Tweets streamed from 01/2012 to 08/2019 and 5M Tweets related to the
COVID-19
pandemic. The general architecture and experimental results of BERTweet can be found in our
paper
:
@inproceedings{bertweet,
title = {{BERTweet: A pre-trained language model for English Tweets}},
author = {Dat Quoc Nguyen and Thanh Vu and Anh Tuan Nguyen},
booktitle = {Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations},
pages = {9--14},
year = {2020}
}
Please CITE
our paper when BERTweet is used to help produce published results or is incorporated into other software.
bertweet-base huggingface.co is an AI model on huggingface.co that provides bertweet-base's model effect (), which can be used instantly with this vinai bertweet-base model. huggingface.co supports a free trial of the bertweet-base model, and also provides paid use of the bertweet-base. Support call bertweet-base model through api, including Node.js, Python, http.
bertweet-base huggingface.co is an online trial and call api platform, which integrates bertweet-base's modeling effects, including api services, and provides a free online trial of bertweet-base, you can try bertweet-base online for free by clicking the link below.
vinai bertweet-base online free url in huggingface.co:
bertweet-base is an open source model from GitHub that offers a free installation service, and any user can find bertweet-base on GitHub to install. At the same time, huggingface.co provides the effect of bertweet-base install, users can directly use bertweet-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.