ModernCamemBERT
is a French language model pretrained on a large corpus of 1T tokens of High-Quality French text. It is the French version of the
ModernBERT
model. ModernCamemBERT was trained using the Masked Language Modeling (MLM) objective with 30% mask rate on 1T tokens on 48 H100 GPUs. The dataset used for training is a combination of French
RedPajama-V2
filtered using heuristic and semantic filtering, French scientific documents from
HALvest
, and the French Wikipedia. Semantic filtering was done by fine-tuning a BERT classifier trained on a document quality dataset automatically labeled by LLama-3 70B.
We also re-use the old
CamemBERTav2
tokenizer. The model was first trained with 1024 context length which was then increased to 8192 tokens later in the pretraining. More details about the training process can be found in the
ModernCamemBERT
paper.
The goal of ModernCamemBERT was to run a controlled study by pretraining ModernBERT on the same dataset as CamemBERTaV2, a DeBERTaV3 French model, isolating the effect of model design. Our results show that the previous model generation remains superior in sample efficiency and overall benchmark performance, with ModernBERT’s primary advantage being faster training and inference speed. However, the new proposed model still provides meaningful architectural improvements compared to earlier models such as the BERT and RoBERTa CamemBERT/v2 model. Additionally, we observe that high-quality pre-training data accelerates convergence but does not significantly improve final performance, suggesting potential benchmark saturation.
We recommend using the ModernCamemBERT model for tasks that require a large context length or efficient inference speed.
Other tasks should still use the CamemBERTaV2 model, which is still the best performing model on most benchmarks.
We release two versions of the model:
almanach/moderncamembert-base
and
almanach/moderncamembert-cv2-base
. The first version is the one trained on the new high-quality 1T token dataset, while the second one is the one trained on the old CamemBERTaV2 dataset. The two models are trained with the same architecture and hyperparameters.
How to use
from transformers import AutoTokenizer, AutoModel, AutoModelForMaskedLM
model = AutoModel.from_pretrained("almanach/moderncamembert-base")
tokenizer = AutoTokenizer.from_pretrained("almanach/moderncamembert-base")
Fine-tuning Results:
Datasets: NER (FTB), the FLUE benchmark (XNLI, CLS, PAWS-X), the French Question Answering Dataset (FQuAD).
moderncamembert-base huggingface.co is an AI model on huggingface.co that provides moderncamembert-base's model effect (), which can be used instantly with this almanach moderncamembert-base model. huggingface.co supports a free trial of the moderncamembert-base model, and also provides paid use of the moderncamembert-base. Support call moderncamembert-base model through api, including Node.js, Python, http.
moderncamembert-base huggingface.co is an online trial and call api platform, which integrates moderncamembert-base's modeling effects, including api services, and provides a free online trial of moderncamembert-base, you can try moderncamembert-base online for free by clicking the link below.
almanach moderncamembert-base online free url in huggingface.co:
moderncamembert-base is an open source model from GitHub that offers a free installation service, and any user can find moderncamembert-base on GitHub to install. At the same time, huggingface.co provides the effect of moderncamembert-base install, users can directly use moderncamembert-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
moderncamembert-base install url in huggingface.co: