MMLW (muszę mieć lepszą wiadomość) are neural text encoders for Polish.
This model is optimized for information retrieval tasks. It can transform queries and passages to 768 dimensional vectors.
The model was developed using a two-step procedure:
The second step involved fine-tuning the obtained models with contrastrive loss on
Polish MS MARCO
training split. In order to improve the efficiency of contrastive training, we used large batch sizes - 1152 for small, 768 for base, and 288 for large models. Fine-tuning was conducted on a cluster of 12 A100 GPUs.
⚠️
2023-12-26:
We have updated the model to a new version with improved results. You can still download the previous version using the
v1
tag:
AutoModel.from_pretrained("sdadas/mmlw-retrieval-e5-base", revision="v1")
⚠️
Usage (Sentence-Transformers)
⚠️ Our dense retrievers require the use of specific prefixes and suffixes when encoding texts. For this model, queries should be prefixed with
"query: "
and passages with
"passage: "
⚠️
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
query_prefix = "query: "
answer_prefix = "passage: "
queries = [query_prefix + "Jak dożyć 100 lat?"]
answers = [
answer_prefix + "Trzeba zdrowo się odżywiać i uprawiać sport.",
answer_prefix + "Trzeba pić alkohol, imprezować i jeździć szybkimi autami.",
answer_prefix + "Gdy trwała kampania politycy zapewniali, że rozprawią się z zakazem niedzielnego handlu."
]
model = SentenceTransformer("sdadas/mmlw-retrieval-e5-base")
queries_emb = model.encode(queries, convert_to_tensor=True, show_progress_bar=False)
answers_emb = model.encode(answers, convert_to_tensor=True, show_progress_bar=False)
best_answer = cos_sim(queries_emb, answers_emb).argmax().item()
print(answers[best_answer])
# Trzeba zdrowo się odżywiać i uprawiać sport.
Evaluation Results
The model achieves
NDCG@10
of
56.09
on the Polish Information Retrieval Benchmark. See
PIRB Leaderboard
for detailed results.
Acknowledgements
This model was trained with the A100 GPU cluster support delivered by the Gdansk University of Technology within the TASK center initiative.
Citation
@article{dadas2024pirb,
title={{PIRB}: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods},
author={Sławomir Dadas and Michał Perełkiewicz and Rafał Poświata},
year={2024},
eprint={2402.13350},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
Runs of sdadas mmlw-retrieval-e5-base on huggingface.co
481
Total runs
0
24-hour runs
-22
3-day runs
-18
7-day runs
258
30-day runs
More Information About mmlw-retrieval-e5-base huggingface.co Model
mmlw-retrieval-e5-base huggingface.co is an AI model on huggingface.co that provides mmlw-retrieval-e5-base's model effect (), which can be used instantly with this sdadas mmlw-retrieval-e5-base model. huggingface.co supports a free trial of the mmlw-retrieval-e5-base model, and also provides paid use of the mmlw-retrieval-e5-base. Support call mmlw-retrieval-e5-base model through api, including Node.js, Python, http.
mmlw-retrieval-e5-base huggingface.co is an online trial and call api platform, which integrates mmlw-retrieval-e5-base's modeling effects, including api services, and provides a free online trial of mmlw-retrieval-e5-base, you can try mmlw-retrieval-e5-base online for free by clicking the link below.
sdadas mmlw-retrieval-e5-base online free url in huggingface.co:
mmlw-retrieval-e5-base is an open source model from GitHub that offers a free installation service, and any user can find mmlw-retrieval-e5-base on GitHub to install. At the same time, huggingface.co provides the effect of mmlw-retrieval-e5-base install, users can directly use mmlw-retrieval-e5-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
mmlw-retrieval-e5-base install url in huggingface.co: