The
IT5
model family represents the first effort in pretraining large-scale sequence-to-sequence transformer models for the Italian language, following the approach adopted by the original
T5 model
.
TThe inference widget is deactivated because the model needs a task-specific seq2seq fine-tuning on a downstream task to be useful in practice.
Model variants
This repository contains the checkpoints for the
base
version of the model. The model was trained for one epoch (1.05M steps) on the
Thoroughly Cleaned Italian mC4 Corpus
(~41B words, ~275GB) using 🤗 Datasets and the
google/t5-v1_1-base
improved configuration. Another version of this model trained on the
OSCAR corpus
is also available under the name
gsarti/it5-base-oscar
. The training procedure is made available
on Github
.
The following table summarizes the parameters for all available models
it5-small
it5-base
(this one)
it5-large
it5-base-oscar
dataset
gsarti/clean_mc4_it
gsarti/clean_mc4_it
gsarti/clean_mc4_it
oscar/unshuffled_deduplicated_it
architecture
google/t5-v1_1-small
google/t5-v1_1-base
google/t5-v1_1-large
t5-base
learning rate
5e-3
5e-3
5e-3
1e-2
steps
1'050'000
1'050'000
2'100'000
258'000
training time
36 hours
101 hours
370 hours
98 hours
ff projection
gated-gelu
gated-gelu
gated-gelu
relu
tie embeds
false
false
false
true
optimizer
adafactor
adafactor
adafactor
adafactor
max seq. length
512
512
512
512
per-device batch size
16
16
8
16
tot. batch size
128
128
64
128
weigth decay
1e-3
1e-3
1e-2
1e-3
validation split size
15K examples
15K examples
15K examples
15K examples
The high training time of
it5-base-oscar
was due to
a bug
in the training script.
For a list of individual model parameters, refer to the
config.json
file in the respective repositories.
Using the models
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained("gsarti/it5-base")
model = AutoModelForSeq2SeqLM.from_pretrained("gsarti/it5-base")
Note: You will need to fine-tune the model on your downstream seq2seq task to use it. See an example
here
.
Flax and Tensorflow versions of the model are also available:
Due to the nature of the web-scraped corpus on which IT5 models were trained, it is likely that their usage could reproduce and amplify pre-existing biases in the data, resulting in potentially harmful content such as racial or gender stereotypes and conspiracist views. For this reason, the study of such biases is explicitly encouraged, and model usage should ideally be restricted to research-oriented and non-user-facing endeavors.
Model curators
For problems or updates on this model, please contact
[email protected]
.
Citation Information
@inproceedings{sarti-nissim-2024-it5-text,
title = "{IT}5: Text-to-text Pretraining for {I}talian Language Understanding and Generation",
author = "Sarti, Gabriele and
Nissim, Malvina",
editor = "Calzolari, Nicoletta and
Kan, Min-Yen and
Hoste, Veronique and
Lenci, Alessandro and
Sakti, Sakriani and
Xue, Nianwen",
booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.lrec-main.823",
pages = "9422--9433",
abstract = "We introduce IT5, the first family of encoder-decoder transformer models pretrained specifically on Italian. We document and perform a thorough cleaning procedure for a large Italian corpus and use it to pretrain four IT5 model sizes. We then introduce the ItaGen benchmark, which includes a broad range of natural language understanding and generation tasks for Italian, and use it to evaluate the performance of IT5 models and multilingual baselines. We find monolingual IT5 models to provide the best scale-to-performance ratio across tested models, consistently outperforming their multilingual counterparts and setting a new state-of-the-art for Italian language generation.",
}
Runs of gsarti it5-base on huggingface.co
76
Total runs
0
24-hour runs
3
3-day runs
10
7-day runs
-23
30-day runs
More Information About it5-base huggingface.co Model
it5-base huggingface.co is an AI model on huggingface.co that provides it5-base's model effect (), which can be used instantly with this gsarti it5-base model. huggingface.co supports a free trial of the it5-base model, and also provides paid use of the it5-base. Support call it5-base model through api, including Node.js, Python, http.
it5-base huggingface.co is an online trial and call api platform, which integrates it5-base's modeling effects, including api services, and provides a free online trial of it5-base, you can try it5-base online for free by clicking the link below.
gsarti it5-base online free url in huggingface.co:
it5-base is an open source model from GitHub that offers a free installation service, and any user can find it5-base on GitHub to install. At the same time, huggingface.co provides the effect of it5-base install, users can directly use it5-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.