mouped / duplicate-detection

huggingface.co
Total runs: 97
24-hour runs: 0
7-day runs: 14
30-day runs: -263
Model's Last Updated: July 27 2026
sentence-similarity

Introduction of duplicate-detection

Model Details of duplicate-detection

SentenceTransformer based on intfloat/multilingual-e5-small

This is a sentence-transformers model finetuned from intfloat/multilingual-e5-small . It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for retrieval.

Model Details
Model Description
  • Model Type: Sentence Transformer
  • Base model: intfloat/multilingual-e5-small
  • Maximum Sequence Length: 512 tokens
  • Output Dimensionality: 384 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text
Model Sources
Full Model Architecture
SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 384, 'pooling_mode': 'mean', 'include_prompt': True})
  (2): Normalize({})
)
Usage
Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'query: What free online course can I take to learn how to be an expert in drawing?',
    'query: What are some best online courses to learn Drawing?',
    'query: Permisi, tagihan WiFi bulan ini belum saya bayar, please advise',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000,  0.7386,  0.0006],
#         [ 0.7386,  1.0000, -0.0875],
#         [ 0.0006, -0.0875,  1.0000]])
Evaluation
Metrics
Binary Classification
Metric Value
cosine_accuracy 0.9811
cosine_accuracy_threshold 0.8151
cosine_f1 0.9781
cosine_f1_threshold 0.8109
cosine_precision 0.9772
cosine_recall 0.9791
cosine_ap 0.9962
cosine_mcc 0.9615
Training Details
Training Dataset
Unnamed Dataset
  • Size: 67,265 training samples
  • Columns: sentence_0 , sentence_1 , and label
  • Approximate statistics based on the first 100 samples:
    sentence_0 sentence_1 label
    type string string float
    modality text text
    details
    • min: 11 tokens
    • mean: 18.45 tokens
    • max: 29 tokens
    • min: 9 tokens
    • mean: 19.12 tokens
    • max: 45 tokens
    • min: 0.0
    • mean: 0.49
    • max: 1.0
  • Samples:
    sentence_0 sentence_1 label
    query: Akses saya ke aplikasi butuh reset password dari kemarin, bisa dibantu ya? query: My 2fa verification code never arrives since last night, can someone assist? 0.0
    query: Mohon segera ditindaklanjuti: internet kantor mati total, bukan lambat, kindly assist query: Selamat pagi, tagihan internet bulan ini lebih mahal dari biasanya, please advise 0.0
    query: The company website won't load since yesterday, can someone assist? query: Need help — Our internal portal keeps throwing a 500 error since this afternoon. 1.0
  • Loss: ContrastiveLoss with these parameters:
    {
        "distance_metric": "SiameseDistanceMetric.COSINE_DISTANCE",
        "margin": 0.5,
        "size_average": true
    }
    
Training Hyperparameters
Non-Default Hyperparameters
  • per_device_train_batch_size : 32
  • num_train_epochs : 2
  • per_device_eval_batch_size : 32
  • multi_dataset_batch_sampler : round_robin
All Hyperparameters
Click to expand
  • per_device_train_batch_size : 32
  • num_train_epochs : 2
  • max_steps : -1
  • learning_rate : 5e-05
  • lr_scheduler_type : linear
  • lr_scheduler_kwargs : None
  • warmup_steps : 0
  • optim : adamw_torch_fused
  • optim_args : None
  • weight_decay : 0.0
  • adam_beta1 : 0.9
  • adam_beta2 : 0.999
  • adam_epsilon : 1e-08
  • optim_target_modules : None
  • gradient_accumulation_steps : 1
  • average_tokens_across_devices : True
  • max_grad_norm : 1
  • label_smoothing_factor : 0.0
  • bf16 : False
  • fp16 : False
  • bf16_full_eval : False
  • fp16_full_eval : False
  • tf32 : None
  • gradient_checkpointing : False
  • gradient_checkpointing_kwargs : None
  • torch_compile : False
  • torch_compile_backend : None
  • torch_compile_mode : None
  • use_liger_kernel : False
  • liger_kernel_config : None
  • use_cache : False
  • neftune_noise_alpha : None
  • torch_empty_cache_steps : None
  • auto_find_batch_size : False
  • log_on_each_node : True
  • logging_nan_inf_filter : True
  • include_num_input_tokens_seen : no
  • log_level : passive
  • log_level_replica : warning
  • disable_tqdm : False
  • project : huggingface
  • trackio_space_id : None
  • trackio_bucket_id : None
  • trackio_static_space_id : None
  • per_device_eval_batch_size : 32
  • prediction_loss_only : True
  • eval_on_start : False
  • eval_do_concat_batches : True
  • eval_use_gather_object : False
  • eval_accumulation_steps : None
  • include_for_metrics : []
  • batch_eval_metrics : False
  • save_only_model : False
  • save_on_each_node : False
  • enable_jit_checkpoint : False
  • push_to_hub : False
  • hub_private_repo : None
  • hub_model_id : None
  • hub_strategy : every_save
  • hub_always_push : False
  • hub_revision : None
  • load_best_model_at_end : False
  • ignore_data_skip : False
  • restore_callback_states_from_checkpoint : False
  • full_determinism : False
  • seed : 42
  • data_seed : None
  • use_cpu : False
  • accelerator_config : {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config : None
  • dataloader_drop_last : False
  • dataloader_num_workers : 0
  • dataloader_pin_memory : True
  • dataloader_persistent_workers : False
  • dataloader_prefetch_factor : None
  • remove_unused_columns : True
  • label_names : None
  • train_sampling_strategy : random
  • length_column_name : length
  • ddp_find_unused_parameters : None
  • ddp_bucket_cap_mb : None
  • ddp_broadcast_buffers : False
  • ddp_static_graph : None
  • ddp_backend : None
  • ddp_timeout : 1800
  • fsdp : None
  • fsdp_config : None
  • deepspeed : None
  • debug : []
  • skip_memory_metrics : True
  • do_predict : False
  • resume_from_checkpoint : None
  • warmup_ratio : None
  • local_rank : -1
  • prompts : None
  • batch_sampler : batch_sampler
  • multi_dataset_batch_sampler : round_robin
  • router_mapping : {}
  • learning_rate_mapping : {}
Training Logs
Epoch Step Training Loss ticket-duplicate-eval_cosine_ap
-1 -1 - 0.6566
0.2378 500 0.0181 0.9753
0.4755 1000 0.0055 0.9911
0.7133 1500 0.0036 0.9937
0.9510 2000 0.0030 0.9943
1.0 2103 - 0.9947
1.1888 2500 0.0026 0.9952
1.4265 3000 0.0024 0.9959
1.6643 3500 0.0022 0.9957
1.9020 4000 0.0021 0.9961
2.0 4206 - 0.9962
-1 -1 - 0.9962
Training Time
  • Training : 43.8 minutes
Framework Versions
  • Python: 3.12.13
  • Sentence Transformers: 5.6.0
  • Transformers: 5.13.1
  • PyTorch: 2.11.0+cu128
  • Accelerate: 1.14.0
  • Datasets: 4.0.0
  • Tokenizers: 0.22.2
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
ContrastiveLoss
@inproceedings{hadsell2006dimensionality,
    author={Hadsell, R. and Chopra, S. and LeCun, Y.},
    booktitle={2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'06)},
    title={Dimensionality Reduction by Learning an Invariant Mapping},
    year={2006},
    volume={2},
    number={},
    pages={1735-1742},
    doi={10.1109/CVPR.2006.100}
}

Runs of mouped duplicate-detection on huggingface.co

97
Total runs
0
24-hour runs
1
3-day runs
14
7-day runs
-263
30-day runs

More Information About duplicate-detection huggingface.co Model

duplicate-detection huggingface.co

duplicate-detection huggingface.co is an AI model on huggingface.co that provides duplicate-detection's model effect (), which can be used instantly with this mouped duplicate-detection model. huggingface.co supports a free trial of the duplicate-detection model, and also provides paid use of the duplicate-detection. Support call duplicate-detection model through api, including Node.js, Python, http.

duplicate-detection huggingface.co Url

https://huggingface.co/mouped/duplicate-detection

mouped duplicate-detection online free

duplicate-detection huggingface.co is an online trial and call api platform, which integrates duplicate-detection's modeling effects, including api services, and provides a free online trial of duplicate-detection, you can try duplicate-detection online for free by clicking the link below.

mouped duplicate-detection online free url in huggingface.co:

https://huggingface.co/mouped/duplicate-detection

duplicate-detection install

duplicate-detection is an open source model from GitHub that offers a free installation service, and any user can find duplicate-detection on GitHub to install. At the same time, huggingface.co provides the effect of duplicate-detection install, users can directly use duplicate-detection installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

duplicate-detection install url in huggingface.co:

https://huggingface.co/mouped/duplicate-detection

Url of duplicate-detection

duplicate-detection huggingface.co Url

Provider of duplicate-detection huggingface.co

mouped
ORGANIZATIONS