A pretrained Korean-specific BERT model developed by Computational Linguistics Lab at Seoul National University.
It is based on our character-level
KR-BERT
model which utilize WordPiece tokenizer.
Here, the model name has a suffix 'MEDIUM' since its training data grew from KR-BERT's original dataset. We have another additional model, KR-BERT-EXPANDED with more extensive training data expanded from those of KR-BERT-MEDIUM, so the suffix 'MEDIUM' is used.
Vocab, Parameters and Data
Mulitlingual BERT
(Google)
KorBERT
(ETRI)
KoBERT
(SKT)
KR-BERT character
KR-BERT-MEDIUM
vocab size
119,547
30,797
8,002
16,424
20,000
parameter size
167,356,416
109,973,391
92,186,880
99,265,066
102,015,010
data size
-
(The Wikipedia data
for 104 languages)
23GB
4.7B morphemes
-
(25M sentences,
233M words)
2.47GB
20M sentences,
233M words
12.37GB
91M sentences,
1.17B words
The training data for this model is expanded from those of KR-BERT, texts from Korean Wikipedia, and news articles, by addition of legal texts crawled from the National Law Information Center and
Korean Comments dataset
. This data expansion is to collect texts from more various domains than those of KR-BERT. The total data size is about 12.37GB, consisting of 91M and 1.17B words.
The user-generated comment dataset is expected to have similar stylistic properties to the task datasets of NSMC and HSD. Such text includes abbreviations, coinages, emoticons, spacing errors, and typos. Therefore, we added the dataset containing such on-line properties to our existing formal data such as news articles and Wikipedia texts to compose the training data for KR-BERT-MEDIUM. Accordingly, KR-BERT-MEDIUM reported better results in sentiment analysis than other models, and the performances improved with the model of the more massive, more various training data.
This model’s vocabulary size is 20,000, whose tokens are trained based on the expanded training data using the WordPiece tokenizer.
KR-BERT-MEDIUM is trained for 2M steps with the maxlen of 128, training batch size of 64, and learning rate of 1e-4, taking 22 hours to train the model using a Google Cloud TPU v3-8.
Models
TensorFlow
BERT tokenizer, character-based model (
download
)
PyTorch
You can import it from Transformers!
# pytorch, transformers
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("snunlp/KR-Medium", do_lower_case=False)
model = AutoModel.from_pretrained("snunlp/KR-Medium")
Requirements
transformers == 4.0.0
tensorflow < 2.0
Downstream tasks
Movie Review Classification on Naver Sentiment Movie Corpus
(NSMC)
More Information About KR-Medium huggingface.co Model
KR-Medium huggingface.co
KR-Medium huggingface.co is an AI model on huggingface.co that provides KR-Medium's model effect (), which can be used instantly with this snunlp KR-Medium model. huggingface.co supports a free trial of the KR-Medium model, and also provides paid use of the KR-Medium. Support call KR-Medium model through api, including Node.js, Python, http.
KR-Medium huggingface.co is an online trial and call api platform, which integrates KR-Medium's modeling effects, including api services, and provides a free online trial of KR-Medium, you can try KR-Medium online for free by clicking the link below.
snunlp KR-Medium online free url in huggingface.co:
KR-Medium is an open source model from GitHub that offers a free installation service, and any user can find KR-Medium on GitHub to install. At the same time, huggingface.co provides the effect of KR-Medium install, users can directly use KR-Medium installed effect in huggingface.co for debugging and trial. It also supports api for free installation.