almanach/camembertv2-base-gsd is a roberta model for token classification. It is trained on the GSD dataset for the task of Part-of-Speech Tagging and Dependency Parsing.
The model achieves an f1 score of on the GSD dataset.
The model is part of the almanach/camembertv2-base family of model finetunes.
Model Details
Model Description
Developed by:
Wissam Antoun (Phd Student at Almanach, Inria-Paris)
The model can be used for token classification tasks in French for Part-of-Speech Tagging and Dependency Parsing.
Bias, Risks, and Limitations
The model may exhibit biases based on the training data. The model may not generalize well to other datasets or tasks. The model may also have limitations in terms of the data it was trained on.
Model trained with the
hopsparser
library on the GSD dataset.
Training Hyperparameters
# Layer dimensionsmlp_input:1024mlp_tag_hidden:16mlp_arc_hidden:512mlp_lab_hidden:128# Lexerslexers:-name:word_embeddingstype:wordsembedding_size:256word_dropout:0.5-name:char_level_embeddingstype:chars_rnnembedding_size:64lstm_output_size:128-name:fasttexttype:fasttext-name:camembertv2_base_p2_17k_last_layertype:bertmodel:/scratch/camembertv2/runs/models/camembertv2-base-bf16/post/ckpt-p2-17000/pt/layers: [11]
subwords_reduction:"mean"# Training hyperparametersencoder_dropout:0.5mlp_dropout:0.5batch_size:8epochs:64lr:base:0.00003schedule:shape:linearwarmup_steps:100
Results
UPOS:
0.98662
LAS:
0.94317
Technical Specifications
Model Architecture and Objective
roberta custom model for token classification.
Citation
BibTeX:
@misc{antoun2024camembert20smarterfrench,
title={CamemBERT 2.0: A Smarter French Language Model Aged to Perfection},
author={Wissam Antoun and Francis Kulumba and Rian Touchent and Éric de la Clergerie and Benoît Sagot and Djamé Seddah},
year={2024},
eprint={2411.08868},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2411.08868},
}
@inproceedings{grobol:hal-03223424,
title = {Analyse en dépendances du français avec des plongements contextualisés},
author = {Grobol, Loïc and Crabbé, Benoît},
url = {https://hal.archives-ouvertes.fr/hal-03223424},
booktitle = {Actes de la 28ème Conférence sur le Traitement Automatique des Langues Naturelles},
eventtitle = {TALN-RÉCITAL 2021},
venue = {Lille, France},
pdf = {https://hal.archives-ouvertes.fr/hal-03223424/file/HOPS_final.pdf},
hal_id = {hal-03223424},
hal_version = {v1},
}
Runs of almanach camembertv2-base-gsd on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About camembertv2-base-gsd huggingface.co Model
camembertv2-base-gsd huggingface.co is an AI model on huggingface.co that provides camembertv2-base-gsd's model effect (), which can be used instantly with this almanach camembertv2-base-gsd model. huggingface.co supports a free trial of the camembertv2-base-gsd model, and also provides paid use of the camembertv2-base-gsd. Support call camembertv2-base-gsd model through api, including Node.js, Python, http.
camembertv2-base-gsd huggingface.co is an online trial and call api platform, which integrates camembertv2-base-gsd's modeling effects, including api services, and provides a free online trial of camembertv2-base-gsd, you can try camembertv2-base-gsd online for free by clicking the link below.
almanach camembertv2-base-gsd online free url in huggingface.co:
camembertv2-base-gsd is an open source model from GitHub that offers a free installation service, and any user can find camembertv2-base-gsd on GitHub to install. At the same time, huggingface.co provides the effect of camembertv2-base-gsd install, users can directly use camembertv2-base-gsd installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
camembertv2-base-gsd install url in huggingface.co: