CAMeL-Lab/text-editing-coda
is a text editing model tailored for grammatical error correction (GEC) in dialectal Arabic (DA).
The model is based on
AraBERTv02
, which we fine-tuned using the
MADAR CODA
corpus.
This model was introduced in our ACL 2025 paper,
Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case Study
, where we refer to it as SWEET (Subword Edit Error Tagger).
It achieved SOTA performance on the MADAR CODA dataset. Details about the training procedure, data preprocessing, and hyperparameters are available in the paper.
The fine-tuning code and associated resources are publicly available on our GitHub repository:
https://github.com/CAMeL-Lab/text-editing
.
Intended uses
To use the
CAMeL-Lab/text-editing-coda
model, you must clone our text editing
GitHub repository
and follow the installation requirements.
We used this
SWEET
model to report results on the MADAR CODA dev and test sets in our
paper
.
How to use
Clone our text editing
GitHub repository
and follow the installation requirements
from transformers import BertTokenizer, BertForTokenClassification
import torch
import torch.nn.functional as F
from gec.tag import rewrite
tokenizer = BertTokenizer.from_pretrained('CAMeL-Lab/text-editing-coda')
model = BertForTokenClassification.from_pretrained('CAMeL-Lab/text-editing-coda')
text = 'أنا بعطيك رقم تلفونو و عنوانو'.split()
tokenized_text = tokenizer(text, return_tensors="pt", is_split_into_words=True)
with torch.no_grad():
logits = model(**tokenized_text).logits
preds = F.softmax(logits.squeeze(), dim=-1)
preds = torch.argmax(preds, dim=-1).cpu().numpy()
edits = [model.config.id2label[p] for p in preds[1:-1]]
assertlen(edits) == len(tokenized_text['input_ids'][0][1:-1])
print(edits) # ['R_[ا]K*', 'K*I_[ا]K', 'K*', 'K*', 'K*', 'K*', 'K*R_[ه]', 'K*', 'MK*', 'R_[ه]']
subwords = tokenizer.convert_ids_to_tokens(tokenized_text['input_ids'][0][1:-1])
output_sent = rewrite(subwords=[subwords], edits=[edits])[0][0]
print(output_sent) # انا باعطيك رقم تلفونه وعنوانه
Citation
@inter{alhafni-habash-2025-enhancing,
title={Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case Study},
author={Bashar Alhafni and Nizar Habash},
year={2025},
eprint={2503.00985},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2503.00985},
}
Runs of CAMeL-Lab text-editing-coda on huggingface.co
59
Total runs
0
24-hour runs
1
3-day runs
16
7-day runs
13
30-day runs
More Information About text-editing-coda huggingface.co Model
text-editing-coda huggingface.co is an AI model on huggingface.co that provides text-editing-coda's model effect (), which can be used instantly with this CAMeL-Lab text-editing-coda model. huggingface.co supports a free trial of the text-editing-coda model, and also provides paid use of the text-editing-coda. Support call text-editing-coda model through api, including Node.js, Python, http.
text-editing-coda huggingface.co is an online trial and call api platform, which integrates text-editing-coda's modeling effects, including api services, and provides a free online trial of text-editing-coda, you can try text-editing-coda online for free by clicking the link below.
CAMeL-Lab text-editing-coda online free url in huggingface.co:
text-editing-coda is an open source model from GitHub that offers a free installation service, and any user can find text-editing-coda on GitHub to install. At the same time, huggingface.co provides the effect of text-editing-coda install, users can directly use text-editing-coda installed effect in huggingface.co for debugging and trial. It also supports api for free installation.