BSC-LT / RoBERTa-ca

huggingface.co
Total runs: 172
24-hour runs: 2
7-day runs: 56
30-day runs: 83
Model's Last Updated: April 10 2026
fill-mask

Introduction of RoBERTa-ca

Model Details of RoBERTa-ca

RoBERTa-ca Model Card

RoBERTa-ca is a new foundational Catalan language model built on the RoBERTa architecture. It uses vocabulary adaptation from mRoBERTa , a method that initializes all weights from mRoBERTa while applying a specialized treatment to the embedding matrix. This treatment carefully handles the differences between the two tokenizers. The model is then continually pretrained using a Catalan-only corpus, consisting of 95GB of high-quality Catalan data.

Technical Description

Technical details of the RoBERTa-ca model.

Description Value
Model Parameters 125M
Tokenizer Type SPM
Vocabulary size 50,304
Precision bfloat16
Context length 512

Training Hyperparemeters

Hyperparameter Value
Pretraining Objective Masked Language Modeling
Learning Rate 3E-05
Learning Rate Scheduler Cosine
Warmup 2425
Optimizer AdamW
Optimizer Hyperparameters AdamW (β1=0.9,β2=0.98,ε =1e-06 )
Optimizer Decay 1E-02
Global Batch Size 1024
Dropout 1E-01
Attention Dropout 1E-01
Activation Function GeLU
EVALUATION: CLUB Benchmark

Model performance in Catalan Language is assessed using the Catalan benchmark CLUB. CLUB (Catalan Language Understanding Benchmark) consists of 6 tasks: Named Entity Recognition (NER), Part-of-Speech Tagging (POS), Semantic Textual Similarity (STS), Text Classification (TC), Textual Entailment (TE), and Question Answering (QA). This benchmark evaluates the model's capabilities in the Catalan language.

The following base foundational models have been considered for the comparison:

Multilingual Foundational Model Number of Parameters Vocab Size Description
BERTa 126M 52K BERTa is a Catalan-specific language model pretrained with Catalan-only data.
BERTinho 109M 30K BERTinho is monolingual BERT model for Galician language.
mBERT 178M 120K Multilingual BERT model pretrained on the top 104 languages with the largest Wikipedia.
mRoBERTa 283M 256K RoBERTa base model pretrained with 35 European languages and a larger vocabulary size.
roberta-base-bne 125M 50K RoBERTa base model pretrained with 570GB of data from web crawlings performed by the National Library of Spain from 2009 to 2019.
RoBERTa-ca 125M 50K RoBERTa-ca is a Catalan-specific language model obtained by using vocabulary adaptation from mRoBERTa.
xlm-roberta-base 279M 250K Foundational RoBERTa model pretrained with CommonCrawl data containing 100 languages.
xlm-roberta-large 561M 250K Foundational RoBERTa model pretrained with CommonCrawl data containing 100 languages.
tasks roberta-base-bne (125M) berta (126M) mBERT (178M) xlm-roberta-base (279M) xlm-roberta-large (561M) roberta-ca (125M) mRoBERTa (283M)
ner (F1) 87.59 89.47 85.89 87.50 89.47 89.70 88.33
pos (F1) 98.64 98.89 98.78 98.91 99.03 99.00 98.98
sts (Person) 74.27 81.39 77.05 75.11 83.49 82.99 79.52
tc (Acc.) 73.86 73.16 72.00 73.05 74.10 72.81 72.41
te (Acc.) 72.27 80.11 75.86 78.27 86.63 82.14 82.38
viquiquad (F1) 82.56 86.74 87.42 86.81 90.35 87.31 87.86
xquad (F1) 60.56 67.38 67.72 68.56 76.08 70.53 69.40
Additional information
Author

The Language Technologies Lab from Barcelona Supercomputing Center.

Contact

For further information, please send an email to [email protected] .

Copyright

Copyright(c) 2025 by Language Technologies Lab, Barcelona Supercomputing Center.

Funding

This work has been promoted and financed by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project ILENIA with reference 2022/TL22/00215337.

Acknowledgements

This project has benefited from the contributions of numerous teams and institutions through data contributions.

In Catalonia, many institutions have been involved in the project. Our thanks to Òmnium Cultural, Parlament de Catalunya, Institut d'Estudis Aranesos, Racó Català, Vilaweb, ACN, Nació Digital, El món and Aquí Berguedà.

At national level, we are especially grateful to our ILENIA project partners: CENID, HiTZ and CiTIUS for their participation. We also extend our genuine gratitude to the Spanish Senate and Congress, Fundación Dialnet, Fundación Elcano and the ‘Instituto Universitario de Sistemas Inteligentes y Aplicaciones Numéricas en Ingeniería (SIANI)’ of the University of Las Palmas de Gran Canaria.

At the international level, we thank the Welsh government, DFKI, Occiglot project, especially Malte Ostendorff, and The Common Crawl Foundation, especially Pedro Ortiz, for their collaboration.

Their valuable efforts have been instrumental in the development of this work.

Disclaimer

Be aware that the model may contain biases or other unintended distortions. When third parties deploy systems or provide services based on this model, or use the model themselves, they bear the responsibility for mitigating any associated risks and ensuring compliance with applicable regulations, including those governing the use of Artificial Intelligence.

The Barcelona Supercomputing Center, as the owner and creator of the model, shall not be held liable for any outcomes resulting from third-party use.

License

Apache License, Version 2.0

Runs of BSC-LT RoBERTa-ca on huggingface.co

172
Total runs
2
24-hour runs
44
3-day runs
56
7-day runs
83
30-day runs

More Information About RoBERTa-ca huggingface.co Model

More RoBERTa-ca license Visit here:

https://choosealicense.com/licenses/apache-2.0

RoBERTa-ca huggingface.co

RoBERTa-ca huggingface.co is an AI model on huggingface.co that provides RoBERTa-ca's model effect (), which can be used instantly with this BSC-LT RoBERTa-ca model. huggingface.co supports a free trial of the RoBERTa-ca model, and also provides paid use of the RoBERTa-ca. Support call RoBERTa-ca model through api, including Node.js, Python, http.

RoBERTa-ca huggingface.co Url

https://huggingface.co/BSC-LT/RoBERTa-ca

BSC-LT RoBERTa-ca online free

RoBERTa-ca huggingface.co is an online trial and call api platform, which integrates RoBERTa-ca's modeling effects, including api services, and provides a free online trial of RoBERTa-ca, you can try RoBERTa-ca online for free by clicking the link below.

BSC-LT RoBERTa-ca online free url in huggingface.co:

https://huggingface.co/BSC-LT/RoBERTa-ca

RoBERTa-ca install

RoBERTa-ca is an open source model from GitHub that offers a free installation service, and any user can find RoBERTa-ca on GitHub to install. At the same time, huggingface.co provides the effect of RoBERTa-ca install, users can directly use RoBERTa-ca installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

RoBERTa-ca install url in huggingface.co:

https://huggingface.co/BSC-LT/RoBERTa-ca

Url of RoBERTa-ca

RoBERTa-ca huggingface.co Url

Provider of RoBERTa-ca huggingface.co

BSC-LT
ORGANIZATIONS

Other API from BSC-LT

huggingface.co

Total runs: 3.4K
Run Growth: 2.7K
Growth Rate: 75.85%
Updated:April 10 2026
huggingface.co

Total runs: 1.8K
Run Growth: -2.1K
Growth Rate: -112.79%
Updated:October 22 2025
huggingface.co

Total runs: 1.1K
Run Growth: 1.0K
Growth Rate: 96.24%
Updated:April 10 2026
huggingface.co

Total runs: 927
Run Growth: -565
Growth Rate: -60.88%
Updated:October 22 2025
huggingface.co

Total runs: 650
Run Growth: 171
Growth Rate: 26.47%
Updated:October 22 2025
huggingface.co

Total runs: 594
Run Growth: 175
Growth Rate: 29.51%
Updated:April 10 2026
huggingface.co

Total runs: 342
Run Growth: 70
Growth Rate: 22.01%
Updated:March 27 2026
huggingface.co

Total runs: 224
Run Growth: -52
Growth Rate: -22.71%
Updated:August 07 2025
huggingface.co

Total runs: 163
Run Growth: -36
Growth Rate: -22.09%
Updated:October 26 2021
huggingface.co

Total runs: 119
Run Growth: 16
Growth Rate: 13.45%
Updated:September 06 2021
huggingface.co

Total runs: 105
Run Growth: 75
Growth Rate: 71.43%
Updated:October 29 2024
huggingface.co

Total runs: 77
Run Growth: -48
Growth Rate: -62.34%
Updated:April 22 2026
huggingface.co

Total runs: 13
Run Growth: 5
Growth Rate: 38.46%
Updated:September 10 2024