First version of Arabic ColBERT.
This model was trained on 100K filtered triplets of the
akhooli/arabic-triplets-1m-curated-sims-len
which has 1 million Arabic (translated) triplets. The dataset was curated from different sources and enriched with similarity score.
More details on the dataset are available in the data card.
If you downloaded the model before July 27th 8 pm (Jerusalem time), please try the current version.
Use the
Ragatouille examples
to learn more,
just replace the pretrained model name and make sure you use Arabic text and split documents for best results.
You can train a better model if you have access to adequate compute (can finetune this model on more data, seed 42 was used tp pick the 100K sample).
Training script
from datasets import load_dataset
from ragatouille import RAGTrainer
sample_size = 100000
ds = load_dataset('akhooli/arabic-triplets-1m-curated-sims-len', split="train", trust_remote_code=True, streaming=True)
# some data processing not in this script (data filtered based on similarity scores) and 100K selected at random
sds = ds.shuffle(seed=42, buffer_size=10_000)
dsf = sds
triplets = []
for item initer(dsf):
triplets.append((item["query"], item["positive"], item["negative"]))
trainer = RAGTrainer(model_name="Arabic-ColBERT-100k", pretrained_model_name="aubmindlab/bert-base-arabertv02", language_code="ar",)
trainer.prepare_training_data(raw_data=triplets, mine_hard_negatives=False)
trainer.train(batch_size=32,
nbits=4, # How many bits will the trained model use when compressing indexes
maxsteps=3125, # Maximum steps hard stop
use_ib_negatives=True, # Use in-batch negative to calculate loss
dim=128, # How many dimensions per embedding. 128 is the default and works well.
learning_rate=1e-5, # Learning rate, small values ([3e-6,3e-5] work best if the base model is BERT-like, 5e-6 is often the sweet spot)
doc_maxlen=256, # Maximum document length. Because of how ColBERT works, smaller chunks (128-256) work very well.
use_relu=False, # Disable ReLU -- doesn't improve performance
warmup_steps="auto", # Defaults to 10%
)
Install
datasets
and
ragatouille
first. Last checkpoint is saved in
.ragatouille/..../colbert
Arabic-ColBERT-100K huggingface.co is an AI model on huggingface.co that provides Arabic-ColBERT-100K's model effect (), which can be used instantly with this akhooli Arabic-ColBERT-100K model. huggingface.co supports a free trial of the Arabic-ColBERT-100K model, and also provides paid use of the Arabic-ColBERT-100K. Support call Arabic-ColBERT-100K model through api, including Node.js, Python, http.
Arabic-ColBERT-100K huggingface.co is an online trial and call api platform, which integrates Arabic-ColBERT-100K's modeling effects, including api services, and provides a free online trial of Arabic-ColBERT-100K, you can try Arabic-ColBERT-100K online for free by clicking the link below.
akhooli Arabic-ColBERT-100K online free url in huggingface.co:
Arabic-ColBERT-100K is an open source model from GitHub that offers a free installation service, and any user can find Arabic-ColBERT-100K on GitHub to install. At the same time, huggingface.co provides the effect of Arabic-ColBERT-100K install, users can directly use Arabic-ColBERT-100K installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Arabic-ColBERT-100K install url in huggingface.co: