This is a
BiomedBERT Hash Nano
model fined-tuned using
sentence-transformers
. It maps sentences & paragraphs to a 128 dimensional dense vector space and can be used for tasks like clustering or semantic search.
The training dataset was generated using a random sample of
PubMed
title-abstract pairs along with similar title pairs. The training workflow was a two step distillation process as follows.
Build a distilled dataset of teacher scores using the
biomedbert-base-reranker
cross-encoder for a separate random sample of title-abstract pairs.
Further fine-tune the model on the distilled dataset using
KLDivLoss
.
Usage (txtai)
This model can be used to build embeddings databases with
txtai
for semantic search and/or as a knowledge source for retrieval augmented generation (RAG).
import txtai
embeddings = txtai.Embeddings(
path="neuml/biomedbert-hash-nano-embeddings",
content=True,
vectors={"trust_remote_code": True}
)
embeddings.index(documents())
# Run a query
embeddings.search("query to run")
from sentence_transformers import SentenceTransformer
sentences = ["This is an example sentence", "Each sentence is converted"]
model = SentenceTransformer("neuml/biomedbert-hash-nano-embeddings", trust_remote_code=True)
embeddings = model.encode(sentences)
print(embeddings)
Usage (Hugging Face Transformers)
The model can also be used directly with Transformers.
from transformers import AutoTokenizer, AutoModel
import torch
# Mean Pooling - Take attention mask into account for correct averagingdefmeanpooling(output, mask):
embeddings = output[0] # First element of model_output contains all token embeddings
mask = mask.unsqueeze(-1).expand(embeddings.size()).float()
return torch.sum(embeddings * mask, 1) / torch.clamp(mask.sum(1), min=1e-9)
# Sentences we want sentence embeddings for
sentences = ['This is an example sentence', 'Each sentence is converted']
# Load model from HuggingFace Hub
tokenizer = AutoTokenizer.from_pretrained("neuml/biomedbert-hash-nano-embeddings", trust_remote_code=True)
model = AutoModel.from_pretrained("neuml/biomedbert-hash-nano-embeddings", trust_remote_code=True)
# Tokenize sentences
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')
# Compute token embeddingswith torch.no_grad():
output = model(**inputs)
# Perform pooling. In this case, mean pooling.
embeddings = meanpooling(output, inputs['attention_mask'])
print("Sentence embeddings:")
print(embeddings)
Evaluation Results
Performance of this model is compared to previously released models trained on medical literature. The most commonly used small embeddings model is also included for comparison.
The following datasets were used to evaluate model performance.
At only 970K parameters this model packs quite a punch. It's competitive with larger models trained on medical literature retaining 98% of the performance of
pubmedbert-base-embeddings
at 0.88% the size. The performance is also better than
all-MiniLM-L6-v2
, a commonly used small model and it's 23x smaller. It also performs much better than the 8M static embeddings model although it is slower given that model is static.
This is a great model to use for smaller datasets and on limited compute / edge devices. Given that it only produces vectors of 128 dimensions, stored vectors also don't need as much space.
biomedbert-hash-nano-embeddings huggingface.co is an AI model on huggingface.co that provides biomedbert-hash-nano-embeddings's model effect (), which can be used instantly with this NeuML biomedbert-hash-nano-embeddings model. huggingface.co supports a free trial of the biomedbert-hash-nano-embeddings model, and also provides paid use of the biomedbert-hash-nano-embeddings. Support call biomedbert-hash-nano-embeddings model through api, including Node.js, Python, http.
biomedbert-hash-nano-embeddings huggingface.co is an online trial and call api platform, which integrates biomedbert-hash-nano-embeddings's modeling effects, including api services, and provides a free online trial of biomedbert-hash-nano-embeddings, you can try biomedbert-hash-nano-embeddings online for free by clicking the link below.
NeuML biomedbert-hash-nano-embeddings online free url in huggingface.co:
biomedbert-hash-nano-embeddings is an open source model from GitHub that offers a free installation service, and any user can find biomedbert-hash-nano-embeddings on GitHub to install. At the same time, huggingface.co provides the effect of biomedbert-hash-nano-embeddings install, users can directly use biomedbert-hash-nano-embeddings installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
biomedbert-hash-nano-embeddings install url in huggingface.co: