impresso-project / OCR-diversely-robust-gte-multilingual-base

huggingface.co
Total runs: 31
24-hour runs: 0
7-day runs: -8
30-day runs: -81
Model's Last Updated: April 16 2026
sentence-similarity

Introduction of OCR-diversely-robust-gte-multilingual-base

Model Details of OCR-diversely-robust-gte-multilingual-base

This is the released model for LREC Paper (insert link)

This is a sentence-transformers model: It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search.

Model Details

This model that was adapted to be more robust to OCR Noise in German and French. This model would be particularly useful for libraries and archives in Central Europe that want to perform semantic search and longitudinal studies within their collections.

This is an Alibaba-NLP/gte-multilingual-base model that was further adapted by (Michail et al., 2025)

Usage (Sentence-Transformers)

Using this model becomes easy when you have sentence-transformers installed:

pip install -U sentence-transformers

Then you can use the model like this:

from sentence_transformers import SentenceTransformer
sentences = ["This is an example sentence", "Each sentence is converted"]

model = SentenceTransformer('impresso-project/halloween_workshop_ocr_robust_preview')
embeddings = model.encode(sentences)
print(embeddings)
Evaluation Results

I will add the model specific evaluation results once the instance is running again.

Training Details
Training Dataset
Contrastive Training

The model was trained with the parameters:

Loss :

sentence_transformers.losses.MultipleNegativesRankingLoss with parameters:

{'scale': 20.0, 'similarity_fct': 'cos_sim'}

Parameters of the fit()-Method:

{
    "epochs": 1,
    "evaluation_steps": 0,
    "evaluator": "NoneType",
    "max_grad_norm": 1,
    "optimizer_class": "<class 'torch.optim.adamw.AdamW'>",
    "optimizer_params": {
        "lr": 2e-05
    },
    "scheduler": "WarmupLinear",
    "steps_per_epoch": null,
    "warmup_steps": 250,
    "weight_decay": 0.01
}
Full Model Architecture
SentenceTransformer(
  (0): Transformer({'max_seq_length': 8192, 'do_lower_case': False}) with Transformer model: NewModel 
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
  (2): Normalize()
)
Citation
BibTeX
Cheap Character Noise for OCR-Robust Multilingual Embeddings (introducing paper)

For details on the adaptation methodology please refer to our paper (published in ACL2025 Findings). If you use our models or methodology, please cite our work.

@inproceedings{michail-etal-2025-cheap,
    title = "Cheap Character Noise for {OCR}-Robust Multilingual Embeddings",
    author = "Michail, Andrianos  and
      Opitz, Juri  and
      Wang, Yining  and
      Meister, Robin  and
      Sennrich, Rico  and
      Clematide, Simon",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.609/",
    doi = "10.18653/v1/2025.findings-acl.609",
    pages = "11705--11716",
    ISBN = "979-8-89176-256-5",
% LREC 2026 citation — to be added
Original Multilingual GTE Model
@inproceedings{zhang2024mgte,
  title={mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval},
  author={Zhang, Xin and Zhang, Yanzhao and Long, Dingkun and Xie, Wen and Dai, Ziqi and Tang, Jialong and Lin, Huan and Yang, Baosong and Xie, Pengjun and Huang, Fei and others},
  booktitle={Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track},
  pages={1393--1412},
  year={2024}
}
About Impresso
Impresso project

Impresso - Media Monitoring of the Past is an interdisciplinary research project that aims to develop and consolidate tools for processing and exploring large collections of media archives across modalities, time, languages and national borders. The first project (2017-2021) was funded by the Swiss National Science Foundation under grant No. CRSII5_173719 and the second project (2023-2027) by the SNSF under grant No. CRSII5_213585 and the Luxembourg National Research Fund under grant No. 17498891.

Copyright

Copyright (C) 2025 The Impresso team.

License

This program is provided as open source under the GNU Affero General Public License v3 or later.


Impresso Project Logo

Runs of impresso-project OCR-diversely-robust-gte-multilingual-base on huggingface.co

31
Total runs
0
24-hour runs
1
3-day runs
-8
7-day runs
-81
30-day runs

More Information About OCR-diversely-robust-gte-multilingual-base huggingface.co Model

More OCR-diversely-robust-gte-multilingual-base license Visit here:

https://choosealicense.com/licenses/agpl-3.0

OCR-diversely-robust-gte-multilingual-base huggingface.co

OCR-diversely-robust-gte-multilingual-base huggingface.co is an AI model on huggingface.co that provides OCR-diversely-robust-gte-multilingual-base's model effect (), which can be used instantly with this impresso-project OCR-diversely-robust-gte-multilingual-base model. huggingface.co supports a free trial of the OCR-diversely-robust-gte-multilingual-base model, and also provides paid use of the OCR-diversely-robust-gte-multilingual-base. Support call OCR-diversely-robust-gte-multilingual-base model through api, including Node.js, Python, http.

OCR-diversely-robust-gte-multilingual-base huggingface.co Url

https://huggingface.co/impresso-project/OCR-diversely-robust-gte-multilingual-base

impresso-project OCR-diversely-robust-gte-multilingual-base online free

OCR-diversely-robust-gte-multilingual-base huggingface.co is an online trial and call api platform, which integrates OCR-diversely-robust-gte-multilingual-base's modeling effects, including api services, and provides a free online trial of OCR-diversely-robust-gte-multilingual-base, you can try OCR-diversely-robust-gte-multilingual-base online for free by clicking the link below.

impresso-project OCR-diversely-robust-gte-multilingual-base online free url in huggingface.co:

https://huggingface.co/impresso-project/OCR-diversely-robust-gte-multilingual-base

OCR-diversely-robust-gte-multilingual-base install

OCR-diversely-robust-gte-multilingual-base is an open source model from GitHub that offers a free installation service, and any user can find OCR-diversely-robust-gte-multilingual-base on GitHub to install. At the same time, huggingface.co provides the effect of OCR-diversely-robust-gte-multilingual-base install, users can directly use OCR-diversely-robust-gte-multilingual-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

OCR-diversely-robust-gte-multilingual-base install url in huggingface.co:

https://huggingface.co/impresso-project/OCR-diversely-robust-gte-multilingual-base

Url of OCR-diversely-robust-gte-multilingual-base

OCR-diversely-robust-gte-multilingual-base huggingface.co Url

Provider of OCR-diversely-robust-gte-multilingual-base huggingface.co

impresso-project
ORGANIZATIONS

Other API from impresso-project