RoBERTa base trained with data from National Library of Spain (BNE)
Model Description
RoBERTa-base-bne is a transformer-based masked language model for the Spanish language. It is based on the
RoBERTa
base model and has been pre-trained using the largest Spanish corpus known to date, with a total of 570GB of clean and deduplicated text processed for this work, compiled from the web crawlings performed by the
National Library of Spain (Biblioteca Nacional de España)
from 2009 to 2019.
To obtain a high-quality training corpus, the corpus has been preprocessed with a pipeline of operations, including among the others, sentence splitting, language detection, filtering of bad-formed sentences and deduplication of repetitive contents. During the process document boundaries are kept. This resulted into 2TB of Spanish clean corpus. Further global deduplication among the corpus is applied, resulting into 570GB of text.
Some of the statistics of the corpus:
Corpora
Number of documents
Number of tokens
Size (GB)
BNE
201,080,084
135,733,450,668
570GB
Tokenization and pre-training
The training corpus has been tokenized using a byte version of Byte-Pair Encoding (BPE) used in the original
RoBERTA
model with a vocabulary size of 50,262 tokens. The RoBERTa-base-bne pre-training consists of a masked language model training that follows the approach employed for the RoBERTa base. The training lasted a total of 48 hours with 16 computing nodes each one with 4 NVIDIA V100 GPUs of 16GB VRAM.
@misc{gutierrezfandino2021spanish,
title={Spanish Language Models},
author={Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Marc Pàmies and Joan Llop-Palao and Joaquín Silveira-Ocampo and Casimiro Pio Carrino and Aitor Gonzalez-Agirre and Carme Armentano-Oller and Carlos Rodriguez-Penagos and Marta Villegas},
year={2021},
eprint={2107.07253},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
Runs of BSC-LT roberta-base-bne on huggingface.co
1.3K
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
-157
30-day runs
More Information About roberta-base-bne huggingface.co Model
roberta-base-bne huggingface.co is an AI model on huggingface.co that provides roberta-base-bne's model effect (), which can be used instantly with this BSC-LT roberta-base-bne model. huggingface.co supports a free trial of the roberta-base-bne model, and also provides paid use of the roberta-base-bne. Support call roberta-base-bne model through api, including Node.js, Python, http.
roberta-base-bne huggingface.co is an online trial and call api platform, which integrates roberta-base-bne's modeling effects, including api services, and provides a free online trial of roberta-base-bne, you can try roberta-base-bne online for free by clicking the link below.
BSC-LT roberta-base-bne online free url in huggingface.co:
roberta-base-bne is an open source model from GitHub that offers a free installation service, and any user can find roberta-base-bne on GitHub to install. At the same time, huggingface.co provides the effect of roberta-base-bne install, users can directly use roberta-base-bne installed effect in huggingface.co for debugging and trial. It also supports api for free installation.