This model can be used to build embeddings databases with
txtai
for semantic search and/or as a knowledge source for retrieval augmented generation (RAG).
import txtai
# Create embeddings
embeddings = txtai.Embeddings(
path="neuml/pubmedbert-base-embeddings-500K",
content=True,
)
embeddings.index(documents())
# Run a query
embeddings.search("query to run")
from sentence_transformers import SentenceTransformer
from sentence_transformers.models import StaticEmbedding
# Initialize a StaticEmbedding module
static = StaticEmbedding.from_model2vec("neuml/pubmedbert-base-embeddings-500K")
model = SentenceTransformer(modules=[static])
sentences = ["This is an example sentence", "Each sentence is converted"]
embeddings = model.encode(sentences)
print(embeddings)
Usage (Model2Vec)
The model can also be used directly with Model2Vec.
from model2vec import StaticModel
# Load a pretrained Model2Vec model
model = StaticModel.from_pretrained("neuml/pubmedbert-base-embeddings-500K")
# Compute text embeddings
sentences = ["This is an example sentence", "Each sentence is converted"]
embeddings = model.encode(sentences)
print(embeddings)
Evaluation Results
The following compares performance of this model against the models previously compared with
PubMedBERT Embeddings
. The following datasets were used to evaluate model performance.
The accuracy trade is notable but it's not terribly significant.
Runtime performance
As another test, let's see how long each model takes to index 120K article abstracts using the following code. All indexing is done with a RTX 3090 GPU.
from datasets import load_dataset
from tqdm import tqdm
from txtai import Embeddings
ds = load_dataset("ccdv/pubmed-summarization", split="train")
embeddings = Embeddings(path="path to model", content=True, backend="numpy")
embeddings.index(tqdm(ds["abstract"]))
Vocabulary pruning doesn't change the runtime performance in this case. But the model is much smaller. Vectors are stored at
int16
precision. This can be beneficial to smaller/lower powered embedded devices and could lead to faster vectorization times.
Training
This model was vocabulary pruned using the following script.
import json
import os
from collections import Counter
from pathlib import Path
import numpy as np
from model2vec import StaticModel
from more_itertools import batched
from sklearn.decomposition import PCA
from tokenlearn.train import collect_means_and_texts
from tokenizers import Tokenizer
from tqdm import tqdm
from txtai.scoring import ScoringFactory
deftokenize(tokenizer):
# Tokenize into dataset
dataset = []
for t in tqdm(batched(texts, 1024)):
encodings = tokenizer.encode_batch_fast(t, add_special_tokens=False)
for e in encodings:
dataset.append((None, e.ids, None))
return dataset
deftokenweights(tokenizer):
dataset = tokenize(tokenizer)
# Build scoring index
scoring = ScoringFactory.create({"method": "bm25", "terms": True})
scoring.index(dataset)
# Calculate mean value of weights array per token
tokens = np.zeros(tokenizer.get_vocab_size())
for x in scoring.idf:
tokens[x] = np.mean(scoring.terms.weights(x)[1])
return tokens
# See PubMedBERT Embeddings 2M model for details on this data
features = "features"
paths = sorted(Path(features).glob("*.json"))
texts, _ = collect_means_and_texts(paths)
# Output model parameters
output = "output path"
params, dims = 500000, 64
path = "pubmedbert-base-embeddings-2M_unweighted"
model = StaticModel.from_pretrained(path)
os.makedirs(output, exist_ok=True)
withopen(f"{path}/tokenizer.json", "r", encoding="utf-8") as f:
config = json.load(f)
# Calculate number of tokens to keep
tokencount = params // model.dim
# Calculate term frequency
freqs = Counter()
for _, ids, _ in tokenize(model.tokenizer):
freqs.update(ids)
# Select top N most common tokens
uids = set(x for x, _ in freqs.most_common(tokencount))
uids = [uid for token, uid in config["model"]["vocab"].items() if uid in uids or token.startswith("[")]
# Get embeddings for uids
model.embedding = model.embedding[uids]
# Select pruned tokens
pairs, index = [], 0for token, uid in config["model"]["vocab"].items():
if uid in uids:
pairs.append((token, index))
index += 1
config["model"]["vocab"] = dict(pairs)
# Write new tokenizerwithopen(f"{output}/tokenizer.json", "w", encoding="utf-8") as f:
json.dump(config, f, indent=2)
model.tokenizer = Tokenizer.from_file(f"{output}/tokenizer.json")
# Re-weight tokens
weights = tokenweights(model.tokenizer)
# Remove NaNs from embedding, if any
embedding = np.nan_to_num(model.embedding)
# Apply PCA
embedding = PCA(n_components=dims).fit_transform(embedding)
# Apply weights
embedding *= weights[:, None]
# Update model embedding and normalize
model.embedding, model.normalize = embedding.astype(np.int16), True
model.save_pretrained(output)
pubmedbert-base-embeddings-500K huggingface.co is an AI model on huggingface.co that provides pubmedbert-base-embeddings-500K's model effect (), which can be used instantly with this NeuML pubmedbert-base-embeddings-500K model. huggingface.co supports a free trial of the pubmedbert-base-embeddings-500K model, and also provides paid use of the pubmedbert-base-embeddings-500K. Support call pubmedbert-base-embeddings-500K model through api, including Node.js, Python, http.
pubmedbert-base-embeddings-500K huggingface.co is an online trial and call api platform, which integrates pubmedbert-base-embeddings-500K's modeling effects, including api services, and provides a free online trial of pubmedbert-base-embeddings-500K, you can try pubmedbert-base-embeddings-500K online for free by clicking the link below.
NeuML pubmedbert-base-embeddings-500K online free url in huggingface.co:
pubmedbert-base-embeddings-500K is an open source model from GitHub that offers a free installation service, and any user can find pubmedbert-base-embeddings-500K on GitHub to install. At the same time, huggingface.co provides the effect of pubmedbert-base-embeddings-500K install, users can directly use pubmedbert-base-embeddings-500K installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
pubmedbert-base-embeddings-500K install url in huggingface.co: