Different from embedding model, reranker uses question and document as input and directly output similarity instead of
embedding.
You can get a relevance score by inputting query and passage to the reranker.
And the score can be mapped to a float value in [0,1] by sigmoid function.
Usage
Using FlagEmbedding
pip install -U FlagEmbedding
Get relevance scores (higher scores indicate more relevance):
from FlagEmbedding import FlagReranker
reranker = FlagReranker('namdp-ptit/ViRanker',
use_fp16=True) # Setting use_fp16 to True speeds up computation with a slight performance degradation
score = reranker.compute_score(['ai là vị vua cuối cùng của việt nam', 'vua bảo đại là vị vua cuối cùng của nước ta'])
print(score) # 13.71875# You can map the scores into 0-1 by set "normalize=True", which will apply sigmoid function to the score
score = reranker.compute_score(['ai là vị vua cuối cùng của việt nam', 'vua bảo đại là vị vua cuối cùng của nước ta'],
normalize=True)
print(score) # 0.99999889840464
scores = reranker.compute_score(
[
['ai là vị vua cuối cùng của việt nam', 'vua bảo đại là vị vua cuối cùng của nước ta'],
['ai là vị vua cuối cùng của việt nam', 'lý nam đế là vị vua đầu tiên của nước ta']
]
)
print(scores) # [13.7265625, -8.53125]# You can map the scores into 0-1 by set "normalize=True", which will apply sigmoid function to the score
scores = reranker.compute_score(
[
['ai là vị vua cuối cùng của việt nam', 'vua bảo đại là vị vua cuối của nước ta'],
['ai là vị vua cuối cùng của việt nam', 'lý nam đế là vị vua đầu tiên của nước ta']
],
normalize=True
)
print(scores) # [0.99999889840464, 0.00019716942196222918]
Using Huggingface transformers
pip install -U transformers
Get relevance scores (higher scores indicate more relevance):
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained('namdp-ptit/ViRanker')
model = AutoModelForSequenceClassification.from_pretrained('namdp-ptit/ViRanker')
model.eval()
pairs = [
['ai là vị vua cuối cùng của việt nam', 'vua bảo đại là vị vua cuối cùng của nước ta'],
['ai là vị vua cuối cùng của việt nam', 'lý nam đế là vị vua đầu tiên của nước ta']
],
with torch.no_grad():
inputs = tokenizer(pairs, padding=True, truncation=True, return_tensors='pt', max_length=512)
scores = model(**inputs, return_dict=True).logits.view(-1, ).float()
print(scores)
Fine-tune
Data Format
Train data should be a json file, where each line is a dict like this:
query
is the query, and
pos
is a list of positive texts,
neg
is a list of negative texts. If you have no negative
texts for a query, you can random sample some from the entire corpus as the negatives.
Besides, for each query in the train data, we used LLMs to generate hard negative for them by asking LLMs to create a
document that is the opposite one of the documents in 'pos'.
Performance
Below is a comparision table of the results we achieved compared to some other pre-trained Cross-Encoders on
the
MS MMarco Passage Reranking - Vi - Dev
dataset.
viranker-mirror huggingface.co is an AI model on huggingface.co that provides viranker-mirror's model effect (), which can be used instantly with this nrl-ai viranker-mirror model. huggingface.co supports a free trial of the viranker-mirror model, and also provides paid use of the viranker-mirror. Support call viranker-mirror model through api, including Node.js, Python, http.
viranker-mirror huggingface.co is an online trial and call api platform, which integrates viranker-mirror's modeling effects, including api services, and provides a free online trial of viranker-mirror, you can try viranker-mirror online for free by clicking the link below.
nrl-ai viranker-mirror online free url in huggingface.co:
viranker-mirror is an open source model from GitHub that offers a free installation service, and any user can find viranker-mirror on GitHub to install. At the same time, huggingface.co provides the effect of viranker-mirror install, users can directly use viranker-mirror installed effect in huggingface.co for debugging and trial. It also supports api for free installation.