Projecte Aina's English-Catalan machine translation model
Model description
This model was trained from scratch using the
Fairseq toolkit
on a combination of English-Catalan datasets,
which after filtering and cleaning comprised 30.023.034 sentence pairs. The model was evaluated on several public datasets comprising different domains.
Intended uses and limitations
You can use this model for machine translation from Catalan to English.
At the time of submission, no measures have been taken to estimate the bias and toxicity embedded in the model.
However, we are well aware that our models may be biased. We intend to conduct research in these areas in the future, and if completed, this model card will be updated.
Training
Training data
The model was trained on a combination of several datasets, including data collected from
Opus
,
HPLT
, an internally created
CA-EN Parallel Corpus
,
and other sources.
Training procedure
Data preparation
All datasets are deduplicated and filtered to remove any sentence pairs with a cosine similarity of less than 0.75.
This is done using sentence embeddings calculated using
LaBSE
.
The filtered datasets are then concatenated to form a final corpus of 30.023.034 parallel sentences and before training
the punctuation is normalized using a modified version of the join-single-file.py script from
SoftCatalà
.
Tokenization
All data is tokenized using sentencepiece, using 50 thousand token sentencepiece model learned from the combination of all filtered training data.
This model is included.
Hyperparameters
The model is based on the Transformer-XLarge proposed by
Subramanian et al.
The following hyperparamenters were set on the Fairseq toolkit:
Hyperparameter
Value
Architecture
transformer_vaswani_wmt_en_de_big
Embedding size
1024
Feedforward size
4096
Number of heads
16
Encoder layers
24
Decoder layers
6
Normalize before attention
True
--share-decoder-input-output-embed
True
--share-all-embeddings
True
Effective batch size
96.000
Optimizer
adam
Adam betas
(0.9, 0.980)
Clip norm
0.0
Learning rate
1e-3
Lr. schedurer
inverse sqrt
Warmup updates
4000
Dropout
0.1
Label smoothing
0.1
The model was trained for a total of 12.500 updates. Weights were saved every 1000 updates and reported results are the average of the last 6 checkpoints.
This work has been promoted and financed by the Generalitat de Catalunya through the
Aina project
.
Disclaimer
Click to expand
The model published in this repository is intended for a generalist purpose and is available to third parties under a permissive Apache License, Version 2.0.
Be aware that the model may have biases and/or any other undesirable distortions.
When third parties deploy or provide systems and/or services to other parties using this model (or any system based on it)
or become users of the model, they should note that it is their responsibility to mitigate the risks arising from its use and,
in any event, to comply with applicable regulations, including regulations regarding the use of Artificial Intelligence.
In no event shall the owner and creator of the model (Barcelona Supercomputing Center)
be liable for any results arising from the use made by third parties.
Runs of projecte-aina aina-translator-ca-en on huggingface.co
76
Total runs
1
24-hour runs
-7
3-day runs
-5
7-day runs
72
30-day runs
More Information About aina-translator-ca-en huggingface.co Model
aina-translator-ca-en huggingface.co is an AI model on huggingface.co that provides aina-translator-ca-en's model effect (), which can be used instantly with this projecte-aina aina-translator-ca-en model. huggingface.co supports a free trial of the aina-translator-ca-en model, and also provides paid use of the aina-translator-ca-en. Support call aina-translator-ca-en model through api, including Node.js, Python, http.
aina-translator-ca-en huggingface.co is an online trial and call api platform, which integrates aina-translator-ca-en's modeling effects, including api services, and provides a free online trial of aina-translator-ca-en, you can try aina-translator-ca-en online for free by clicking the link below.
projecte-aina aina-translator-ca-en online free url in huggingface.co:
aina-translator-ca-en is an open source model from GitHub that offers a free installation service, and any user can find aina-translator-ca-en on GitHub to install. At the same time, huggingface.co provides the effect of aina-translator-ca-en install, users can directly use aina-translator-ca-en installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
aina-translator-ca-en install url in huggingface.co: