This is a
PubMedBERT-base
model fined-tuned using
sentence-transformers
. It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. The training dataset was generated using a random sample of
PubMed
title-abstract pairs along with similar title pairs.
PubMedBERT Embeddings produces higher quality embeddings than generalized models for medical literature. Further fine-tuning for a medical subdomain will result in even better performance.
Usage (txtai)
This model can be used to build embeddings databases with
txtai
for semantic search and/or as a knowledge source for retrieval augmented generation (RAG).
import txtai
embeddings = txtai.Embeddings(path="neuml/pubmedbert-base-embeddings", content=True)
embeddings.index(documents())
# Run a query
embeddings.search("query to run")
from sentence_transformers import SentenceTransformer
sentences = ["This is an example sentence", "Each sentence is converted"]
model = SentenceTransformer("neuml/pubmedbert-base-embeddings")
embeddings = model.encode(sentences)
print(embeddings)
Usage (Hugging Face Transformers)
The model can also be used directly with Transformers.
from transformers import AutoTokenizer, AutoModel
import torch
# Mean Pooling - Take attention mask into account for correct averagingdefmeanpooling(output, mask):
embeddings = output[0] # First element of model_output contains all token embeddings
mask = mask.unsqueeze(-1).expand(embeddings.size()).float()
return torch.sum(embeddings * mask, 1) / torch.clamp(mask.sum(1), min=1e-9)
# Sentences we want sentence embeddings for
sentences = ['This is an example sentence', 'Each sentence is converted']
# Load model from HuggingFace Hub
tokenizer = AutoTokenizer.from_pretrained("neuml/pubmedbert-base-embeddings")
model = AutoModel.from_pretrained("neuml/pubmedbert-base-embeddings")
# Tokenize sentences
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')
# Compute token embeddingswith torch.no_grad():
output = model(**inputs)
# Perform pooling. In this case, mean pooling.
embeddings = meanpooling(output, inputs['attention_mask'])
print("Sentence embeddings:")
print(embeddings)
Evaluation Results
Performance of this model compared to the top base models on the
MTEB leaderboard
is shown below. A popular smaller model was also evaluated along with the most downloaded PubMed similarity model on the Hugging Face Hub.
The following datasets were used to evaluate model performance.
pubmedbert-base-embeddings huggingface.co is an AI model on huggingface.co that provides pubmedbert-base-embeddings's model effect (), which can be used instantly with this NeuML pubmedbert-base-embeddings model. huggingface.co supports a free trial of the pubmedbert-base-embeddings model, and also provides paid use of the pubmedbert-base-embeddings. Support call pubmedbert-base-embeddings model through api, including Node.js, Python, http.
pubmedbert-base-embeddings huggingface.co is an online trial and call api platform, which integrates pubmedbert-base-embeddings's modeling effects, including api services, and provides a free online trial of pubmedbert-base-embeddings, you can try pubmedbert-base-embeddings online for free by clicking the link below.
NeuML pubmedbert-base-embeddings online free url in huggingface.co:
pubmedbert-base-embeddings is an open source model from GitHub that offers a free installation service, and any user can find pubmedbert-base-embeddings on GitHub to install. At the same time, huggingface.co provides the effect of pubmedbert-base-embeddings install, users can directly use pubmedbert-base-embeddings installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
pubmedbert-base-embeddings install url in huggingface.co: