An extended BERT tokenizer
built upon
bert-base-multilingual-cased
, optimized for Dhivehi (Divehi/Thaana script).
Overview
This tokenizer preserves all English and multilingual coverage of the base model, while adding the
top 100,000 high-frequency Dhivehi tokens
extracted from a large corpus. It ensures robust tokenization for both English and Dhivehi with no regressions.
How It Was Built
Base model
:
bert-base-multilingual-cased
Corpus
: ~16.7M lines of cleaned Dhivehi text (one sentence per line)
Processing steps
:
Count word-level tokens in batches of 100,000 lines
Filter out rare tokens (frequency < 5)
Select top 100,000 tokens by frequency
Extend vocab using
tokenizer.add_tokens([...])
Usage Example
from transformers import BertTokenizer
tokenizer = BertTokenizer.from_pretrained("alakxender/bert-dhivehi-tokenizer-extended")
# Tokenization test
text_dv = "ޖެންޑާގެ ސްޓޭޓް އަޒްރާ"print(tokenizer.tokenize(text_dv))
English tokenization unchanged
Dhivehi text now tokenizes into meaningful word-level tokens with zero
[UNK]
Intended Uses
Downstream
NER
,
QA
,
classification
, or
masking
tasks in Dhivehi
Fine-tuning of
bert-base-multilingual-cased
model with extended vocabulary
Specially useful for projects in Maldivian/Thaana script
Limitations
Vocabulary capped at
100,000 new tokens
; may miss very rare words
Tokenization remains
word-level
, not subword-based—better performance may require further tokenizer retraining
Files
vocab.txt
,
tokenizer_config.json
,
special_tokens_map.json
→ Standard tokenizer files
Note
: No
tokenizer.json
; requires
BertTokenizer
, not
Fast
, due to manual extension
Runs of alakxender bert-dhivehi-tokenizer-extended on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About bert-dhivehi-tokenizer-extended huggingface.co Model
More bert-dhivehi-tokenizer-extended license Visit here:
bert-dhivehi-tokenizer-extended huggingface.co is an AI model on huggingface.co that provides bert-dhivehi-tokenizer-extended's model effect (), which can be used instantly with this alakxender bert-dhivehi-tokenizer-extended model. huggingface.co supports a free trial of the bert-dhivehi-tokenizer-extended model, and also provides paid use of the bert-dhivehi-tokenizer-extended. Support call bert-dhivehi-tokenizer-extended model through api, including Node.js, Python, http.
bert-dhivehi-tokenizer-extended huggingface.co is an online trial and call api platform, which integrates bert-dhivehi-tokenizer-extended's modeling effects, including api services, and provides a free online trial of bert-dhivehi-tokenizer-extended, you can try bert-dhivehi-tokenizer-extended online for free by clicking the link below.
alakxender bert-dhivehi-tokenizer-extended online free url in huggingface.co:
bert-dhivehi-tokenizer-extended is an open source model from GitHub that offers a free installation service, and any user can find bert-dhivehi-tokenizer-extended on GitHub to install. At the same time, huggingface.co provides the effect of bert-dhivehi-tokenizer-extended install, users can directly use bert-dhivehi-tokenizer-extended installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
bert-dhivehi-tokenizer-extended install url in huggingface.co: