Introduction of OCR-diversely-robust-gte-multilingual-base
Model Details of OCR-diversely-robust-gte-multilingual-base
This is the released model for LREC Paper (insert link)
This is a
sentence-transformers
model: It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search.
Model Details
This model that was adapted to be more robust to OCR Noise in German and French. This model would be particularly useful for libraries and archives in Central Europe that want to perform semantic search and longitudinal studies within their collections.
from sentence_transformers import SentenceTransformer
sentences = ["This is an example sentence", "Each sentence is converted"]
model = SentenceTransformer('impresso-project/halloween_workshop_ocr_robust_preview')
embeddings = model.encode(sentences)
print(embeddings)
Evaluation Results
I will add the model specific evaluation results once the instance is running again.
Training Details
Training Dataset
Contrastive Training
The model was trained with the parameters:
Loss
:
sentence_transformers.losses.MultipleNegativesRankingLoss
with parameters:
Cheap Character Noise for OCR-Robust Multilingual Embeddings (introducing paper)
For details on the adaptation methodology please refer to our paper (published in ACL2025 Findings). If you use our models or methodology, please cite our work.
@inproceedings{michail-etal-2025-cheap,
title = "Cheap Character Noise for {OCR}-Robust Multilingual Embeddings",
author = "Michail, Andrianos and
Opitz, Juri and
Wang, Yining and
Meister, Robin and
Sennrich, Rico and
Clematide, Simon",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.findings-acl.609/",
doi = "10.18653/v1/2025.findings-acl.609",
pages = "11705--11716",
ISBN = "979-8-89176-256-5",
% LREC 2026 citation — to be added
Original Multilingual GTE Model
@inproceedings{zhang2024mgte,
title={mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval},
author={Zhang, Xin and Zhang, Yanzhao and Long, Dingkun and Xie, Wen and Dai, Ziqi and Tang, Jialong and Lin, Huan and Yang, Baosong and Xie, Pengjun and Huang, Fei and others},
booktitle={Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track},
pages={1393--1412},
year={2024}
}
About Impresso
Impresso project
Impresso - Media Monitoring of the Past
is an interdisciplinary research project that aims to develop and consolidate tools for processing and exploring large collections of media archives across modalities, time, languages and national borders. The first project (2017-2021) was funded by the Swiss National Science Foundation under grant No.
CRSII5_173719
and the second project (2023-2027) by the SNSF under grant No.
CRSII5_213585
and the Luxembourg National Research Fund under grant No. 17498891.
OCR-diversely-robust-gte-multilingual-base huggingface.co is an AI model on huggingface.co that provides OCR-diversely-robust-gte-multilingual-base's model effect (), which can be used instantly with this impresso-project OCR-diversely-robust-gte-multilingual-base model. huggingface.co supports a free trial of the OCR-diversely-robust-gte-multilingual-base model, and also provides paid use of the OCR-diversely-robust-gte-multilingual-base. Support call OCR-diversely-robust-gte-multilingual-base model through api, including Node.js, Python, http.
OCR-diversely-robust-gte-multilingual-base huggingface.co is an online trial and call api platform, which integrates OCR-diversely-robust-gte-multilingual-base's modeling effects, including api services, and provides a free online trial of OCR-diversely-robust-gte-multilingual-base, you can try OCR-diversely-robust-gte-multilingual-base online for free by clicking the link below.
impresso-project OCR-diversely-robust-gte-multilingual-base online free url in huggingface.co:
OCR-diversely-robust-gte-multilingual-base is an open source model from GitHub that offers a free installation service, and any user can find OCR-diversely-robust-gte-multilingual-base on GitHub to install. At the same time, huggingface.co provides the effect of OCR-diversely-robust-gte-multilingual-base install, users can directly use OCR-diversely-robust-gte-multilingual-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
OCR-diversely-robust-gte-multilingual-base install url in huggingface.co: