XLM-RoBERTa-large with continued pretraining on speech transcripts of the
Visual History Archive
.
Part 1 of the used training data is ASR'd using domain-specific Wav2Vec 2.0 and general-domain Zipformer models deployed at
UWebASR
.
Part 2 is machine translated from Part 1 using
MADLAD-400-3B-MT
.
Training Data
ASR data:
cs, de, en, hu, nl, pl
MT data:
cs, da, de, en, hu, nl, pl
Total tokens:
4.9B
Training tokens:
4.4B
Test tokens:
490M
The same documents are used in all 7 languages, but their proportions in terms of the number of tokens might differ.
A random split of 10% is used as a test dataset, preserving the language proportions of the training data. The test set has been masked with 15% probability.
The data preprocessing (reading, tokenization, concatenation, splitting, and masking of the test dataset) takes around 2.5 hours per language using 8 CPUs.
Training Details
Parameters are mostly replicated from [1] Appendix B:
AdamW with eps=1e-6, beta1=0.9, beta2=0.98, weight decay=0.01, learning rate=1e-5 with linear schedule and linear warmup for 6% of the first training steps.
Trained with dynamic masking on 4 L40s with per-device batch size 8, using 64 gradient accumulation steps for an effective batch size of 2048, for 1 epoch (34k steps) on an MLM objective.
Main differences from XLM-RoBERTa-large:
AdamW instead of Adam, effective batch size 2048 instead of 8192, and 34k steps instead of 500k due to a smaller dataset.
Smaller learning rate, since greater ones lead to overfitting.
This somewhat aligns with [2] and [3], who continue the pretraining on small data.
The training takes around 24 hours but can be significantly reduced with more GPUs.
Evaluation
Since the model sees translations of evaluation samples during the training, an additional domain-specific dataset has been prepared for unbiased evaluation.
For this dataset, sentences have been extracted from the EHRI-NER dataset based on
EHRI Online Editions
in 9 languages, not including Danish [4].
It is split into two evaluation datasets EHRI-6 (714k tokens) and EHRI-9 (877k tokens), the latter one including 3 unseen languages (French, Slovak, Yiddish).
Improvements from the XLM-RoBERTa-large checkpoint.
The 490M test set is split from the dataset used to train this model and has a greater proportion of machine translations than the 42M test set.
Perplexity per language in the EHRI data, number of tokens given in parentheses:
xlm-roberta-malach huggingface.co is an AI model on huggingface.co that provides xlm-roberta-malach's model effect (), which can be used instantly with this ufal xlm-roberta-malach model. huggingface.co supports a free trial of the xlm-roberta-malach model, and also provides paid use of the xlm-roberta-malach. Support call xlm-roberta-malach model through api, including Node.js, Python, http.
xlm-roberta-malach huggingface.co is an online trial and call api platform, which integrates xlm-roberta-malach's modeling effects, including api services, and provides a free online trial of xlm-roberta-malach, you can try xlm-roberta-malach online for free by clicking the link below.
ufal xlm-roberta-malach online free url in huggingface.co:
xlm-roberta-malach is an open source model from GitHub that offers a free installation service, and any user can find xlm-roberta-malach on GitHub to install. At the same time, huggingface.co provides the effect of xlm-roberta-malach install, users can directly use xlm-roberta-malach installed effect in huggingface.co for debugging and trial. It also supports api for free installation.