This is a release of Korean-specific, small-scale BERT models with comparable or better performances developed by Computational Linguistics Lab at Seoul National University, referenced in
KR-BERT: A Small-Scale Korean-Specific Language Model
.
Vocab, Parameters and Data
Mulitlingual BERT
(Google)
KorBERT
(ETRI)
KoBERT
(SKT)
KR-BERT character
KR-BERT sub-character
vocab size
119,547
30,797
8,002
16,424
12,367
parameter size
167,356,416
109,973,391
92,186,880
99,265,066
96,145,233
data size
-
(The Wikipedia data
for 104 languages)
23GB
4.7B morphemes
-
(25M sentences,
233M words)
2.47GB
20M sentences,
233M words
2.47GB
20M sentences,
233M words
Model
Masked LM Accuracy
KoBERT
0.750
KR-BERT character BidirectionalWordPiece
0.779
KR-BERT sub-character BidirectionalWordPiece
0.769
Sub-character
Korean text is basically represented with Hangul syllable characters, which can be decomposed into sub-characters, or graphemes. To accommodate such characteristics, we trained a new vocabulary and BERT model on two different representations of a corpus: syllable characters and sub-characters.
In case of using our sub-character model, you should preprocess your data with the code below.
import torch
from transformers import BertConfig, BertModel, BertForPreTraining, BertTokenizer
from unicodedata import normalize
tokenizer_krbert = BertTokenizer.from_pretrained('/path/to/vocab_file.txt', do_lower_case=False)
# convert a string into sub-chardefto_subchar(string):
return normalize('NFKD', string)
sentence = '토크나이저 예시입니다.'print(tokenizer_krbert.tokenize(to_subchar(sentence)))
Tokenization
BidirectionalWordPiece Tokenizer
We use the BidirectionalWordPiece model to reduce search costs while maintaining the possibility of choice. This model applies BPE in both forward and backward directions to obtain two candidates and chooses the one that has a higher frequency.
If you want to use the sub-character version of our models, let the
subchar
argument be
True
.
And you can use the original BERT WordPiece tokenizer by entering
bert
for the
tokenizer
argument, and if you use
ranked
you can use our BidirectionalWordPiece tokenizer.
tensorflow: After downloading our pretrained models, put them in a
models
directory in the
krbert_tensorflow
directory.
pytorch: After downloading our pretrained models, put them in a
pretrained
directory in the
krbert_pytorch
directory.
If you use these models, please cite the following paper:
@article{lee2020krbert,
title={KR-BERT: A Small-Scale Korean-Specific Language Model},
author={Sangah Lee and Hansol Jang and Yunmee Baik and Suzi Park and Hyopil Shin},
year={2020},
journal={ArXiv},
volume={abs/2008.03979}
}
Runs of snunlp KR-BERT-char16424 on huggingface.co
329
Total runs
-37
24-hour runs
-20
3-day runs
2
7-day runs
-149
30-day runs
More Information About KR-BERT-char16424 huggingface.co Model
KR-BERT-char16424 huggingface.co
KR-BERT-char16424 huggingface.co is an AI model on huggingface.co that provides KR-BERT-char16424's model effect (), which can be used instantly with this snunlp KR-BERT-char16424 model. huggingface.co supports a free trial of the KR-BERT-char16424 model, and also provides paid use of the KR-BERT-char16424. Support call KR-BERT-char16424 model through api, including Node.js, Python, http.
KR-BERT-char16424 huggingface.co is an online trial and call api platform, which integrates KR-BERT-char16424's modeling effects, including api services, and provides a free online trial of KR-BERT-char16424, you can try KR-BERT-char16424 online for free by clicking the link below.
snunlp KR-BERT-char16424 online free url in huggingface.co:
KR-BERT-char16424 is an open source model from GitHub that offers a free installation service, and any user can find KR-BERT-char16424 on GitHub to install. At the same time, huggingface.co provides the effect of KR-BERT-char16424 install, users can directly use KR-BERT-char16424 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.