ufal / xlm-roberta-malach

huggingface.co
Total runs: 88
24-hour runs: 0
7-day runs: 1
30-day runs: 1
Model's Last Updated: February 23 2026
fill-mask

Introduction of xlm-roberta-malach

Model Details of xlm-roberta-malach

XLM-RoBERTa-malach

XLM-RoBERTa-large with continued pretraining on speech transcripts of the Visual History Archive .
Part 1 of the used training data is ASR'd using domain-specific Wav2Vec 2.0 and general-domain Zipformer models deployed at UWebASR .
Part 2 is machine translated from Part 1 using MADLAD-400-3B-MT .

Training Data

ASR data: cs, de, en, hu, nl, pl
MT data: cs, da, de, en, hu, nl, pl

Total tokens: 4.9B
Training tokens: 4.4B
Test tokens: 490M

The same documents are used in all 7 languages, but their proportions in terms of the number of tokens might differ. A random split of 10% is used as a test dataset, preserving the language proportions of the training data. The test set has been masked with 15% probability.

The data preprocessing (reading, tokenization, concatenation, splitting, and masking of the test dataset) takes around 2.5 hours per language using 8 CPUs.

Training Details

Parameters are mostly replicated from [1] Appendix B:
AdamW with eps=1e-6, beta1=0.9, beta2=0.98, weight decay=0.01, learning rate=1e-5 with linear schedule and linear warmup for 6% of the first training steps. Trained with dynamic masking on 4 L40s with per-device batch size 8, using 64 gradient accumulation steps for an effective batch size of 2048, for 1 epoch (34k steps) on an MLM objective.

Main differences from XLM-RoBERTa-large:
AdamW instead of Adam, effective batch size 2048 instead of 8192, and 34k steps instead of 500k due to a smaller dataset. Smaller learning rate, since greater ones lead to overfitting. This somewhat aligns with [2] and [3], who continue the pretraining on small data.

The training takes around 24 hours but can be significantly reduced with more GPUs.

Evaluation

Since the model sees translations of evaluation samples during the training, an additional domain-specific dataset has been prepared for unbiased evaluation. For this dataset, sentences have been extracted from the EHRI-NER dataset based on EHRI Online Editions in 9 languages, not including Danish [4]. It is split into two evaluation datasets EHRI-6 (714k tokens) and EHRI-9 (877k tokens), the latter one including 3 unseen languages (French, Slovak, Yiddish).

Perplexity (VHA): 2.5257 -> 1.9064
Perplexity (EHRI-6): 3.1897 -> 2.9683
Perplexity (EHRI-9): 3.1806 -> 3.0340

Improvements from the XLM-RoBERTa-large checkpoint. The 490M test set is split from the dataset used to train this model and has a greater proportion of machine translations than the 42M test set.

Perplexity per language in the EHRI data, number of tokens given in parentheses:

Model cs (195k) de (356k) en (81k) fr (3.5k) hu (45k) nl (2.5k) pl (34k) sk (6k) yi (151k)
XLM-RoBERTa-large 3.1553 3.4038 3.0588 2.0579 2.8928 2.9133 2.5284 2.6245 4.0217
XLM-RoBERTa-malach 2.8023 3.1704 2.9022 2.0254 2.8285 2.8797 2.4003 2.5914 4.0910
References

[1] RoBERTa: A Robustly Optimized BERT Pretraining Approach
[2] Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks
[3] The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary Proceedings
[4] Repurposing Holocaust-Related Digital Scholarly Editions to Develop Multilingual Domain-Specific Named Entity Recognition Tools

Runs of ufal xlm-roberta-malach on huggingface.co

88
Total runs
0
24-hour runs
1
3-day runs
1
7-day runs
1
30-day runs

More Information About xlm-roberta-malach huggingface.co Model

More xlm-roberta-malach license Visit here:

https://choosealicense.com/licenses/mit

xlm-roberta-malach huggingface.co

xlm-roberta-malach huggingface.co is an AI model on huggingface.co that provides xlm-roberta-malach's model effect (), which can be used instantly with this ufal xlm-roberta-malach model. huggingface.co supports a free trial of the xlm-roberta-malach model, and also provides paid use of the xlm-roberta-malach. Support call xlm-roberta-malach model through api, including Node.js, Python, http.

xlm-roberta-malach huggingface.co Url

https://huggingface.co/ufal/xlm-roberta-malach

ufal xlm-roberta-malach online free

xlm-roberta-malach huggingface.co is an online trial and call api platform, which integrates xlm-roberta-malach's modeling effects, including api services, and provides a free online trial of xlm-roberta-malach, you can try xlm-roberta-malach online for free by clicking the link below.

ufal xlm-roberta-malach online free url in huggingface.co:

https://huggingface.co/ufal/xlm-roberta-malach

xlm-roberta-malach install

xlm-roberta-malach is an open source model from GitHub that offers a free installation service, and any user can find xlm-roberta-malach on GitHub to install. At the same time, huggingface.co provides the effect of xlm-roberta-malach install, users can directly use xlm-roberta-malach installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

xlm-roberta-malach install url in huggingface.co:

https://huggingface.co/ufal/xlm-roberta-malach

Url of xlm-roberta-malach

xlm-roberta-malach huggingface.co Url

Provider of xlm-roberta-malach huggingface.co

ufal
ORGANIZATIONS

Other API from ufal

huggingface.co

Total runs: 2.5K
Run Growth: 625
Growth Rate: 24.65%
Updated:September 30 2024
huggingface.co

Total runs: 32
Run Growth: 23
Growth Rate: 71.88%
Updated:April 04 2024
huggingface.co

Total runs: 21
Run Growth: 2
Growth Rate: 9.52%
Updated:April 04 2024
huggingface.co

Total runs: 19
Run Growth: 2
Growth Rate: 10.53%
Updated:April 04 2024
huggingface.co

Total runs: 16
Run Growth: 8
Growth Rate: 50.00%
Updated:April 04 2024
huggingface.co

Total runs: 10
Run Growth: 9
Growth Rate: 90.00%
Updated:September 23 2025