GPT2-large trained with data from National Library of Spain (BNE)
Model Description
GPT2-large-bne is a transformer-based model for the Spanish language. It is based on the
GPT-2
model and has been pre-trained using the largest Spanish corpus known to date, with a total of 570GB of clean and deduplicated text processed for this work, compiled from the web crawlings performed by the
National Library of Spain (Biblioteca Nacional de España)
from 2009 to 2019.
To obtain a high-quality training corpus, the corpus has been preprocessed with a pipeline of operations, including among the others, sentence splitting, language detection, filtering of bad-formed sentences and deduplication of repetitive contents. During the process document boundaries are kept. This resulted into 2TB of Spanish clean corpus. Further global deduplication among the corpus is applied, resulting into 570GB of text.
Some of the statistics of the corpus:
Corpora
Number of documents
Number of tokens
Size (GB)
BNE
201,080,084
135,733,450,668
570GB
Tokenization and pre-training
The training corpus has been tokenized using a byte version of Byte-Pair Encoding (BPE) used in the original
GPT-2
model with a vocabulary size of 50,262 tokens. The GPT2-large-bne pre-training consists of an autoregressive language model training that follows the approach of the GPT-2. The training lasted a total of 10 days with 32 computing nodes each one with 4 NVIDIA V100 GPUs of 16GB VRAM.
@misc{gutierrezfandino2021spanish,
title={Spanish Language Models},
author={Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Marc Pàmies and Joan Llop-Palao and Joaquín Silveira-Ocampo and Casimiro Pio Carrino and Aitor Gonzalez-Agirre and Carme Armentano-Oller and Carlos Rodriguez-Penagos and Marta Villegas},
year={2021},
eprint={2107.07253},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
Runs of BSC-LT gpt2-large-bne on huggingface.co
119
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
16
30-day runs
More Information About gpt2-large-bne huggingface.co Model
gpt2-large-bne huggingface.co is an AI model on huggingface.co that provides gpt2-large-bne's model effect (), which can be used instantly with this BSC-LT gpt2-large-bne model. huggingface.co supports a free trial of the gpt2-large-bne model, and also provides paid use of the gpt2-large-bne. Support call gpt2-large-bne model through api, including Node.js, Python, http.
gpt2-large-bne huggingface.co is an online trial and call api platform, which integrates gpt2-large-bne's modeling effects, including api services, and provides a free online trial of gpt2-large-bne, you can try gpt2-large-bne online for free by clicking the link below.
BSC-LT gpt2-large-bne online free url in huggingface.co:
gpt2-large-bne is an open source model from GitHub that offers a free installation service, and any user can find gpt2-large-bne on GitHub to install. At the same time, huggingface.co provides the effect of gpt2-large-bne install, users can directly use gpt2-large-bne installed effect in huggingface.co for debugging and trial. It also supports api for free installation.