This is a
BERT Hash Nano
model fined-tuned using
sentence-transformers
. It maps sentences & paragraphs to a 128-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
This model is an alternative to
MUVERA fixed-dimensional encoding
with ColBERT models. MUVERA encoding enables encoding the multi-vector outputs of ColBERT into single dense vector outputs. While this is a great step, the main issue with MUVERA is that it tends to need wide vectors to be effective (5K - 10K dimensional vectors).
bert-hash-nano-embeddings
outputs 128-dimensional vectors.
Build a distilled dataset of teacher scores using the
mixedbread-ai/mxbai-rerank-xsmall-v1
cross-encoder for a random sample of the training dataset mentioned above.
Further fine-tune the model on the distilled dataset using
KLDivLoss
.
Usage (txtai)
This model can be used to build embeddings databases with
txtai
for semantic search and/or as a knowledge source for retrieval augmented generation (RAG).
import txtai
embeddings = txtai.Embeddings(
path="neuml/bert-hash-nano-embeddings",
content=True,
vectors={"trust_remote_code": True}
)
embeddings.index(documents())
# Run a query
embeddings.search("query to run")
from sentence_transformers import SentenceTransformer
sentences = ["This is an example sentence", "Each sentence is converted"]
model = SentenceTransformer("neuml/bert-hash-nano-embeddings", trust_remote_code=True)
embeddings = model.encode(sentences)
print(embeddings)
Usage (Hugging Face Transformers)
The model can also be used directly with Transformers.
from transformers import AutoTokenizer, AutoModel
import torch
# Mean Pooling - Take attention mask into account for correct averagingdefmeanpooling(output, mask):
embeddings = output[0] # First element of model_output contains all token embeddings
mask = mask.unsqueeze(-1).expand(embeddings.size()).float()
return torch.sum(embeddings * mask, 1) / torch.clamp(mask.sum(1), min=1e-9)
# Sentences we want sentence embeddings for
sentences = ['This is an example sentence', 'Each sentence is converted']
# Load model from HuggingFace Hub
tokenizer = AutoTokenizer.from_pretrained("neuml/bert-hash-nano-embeddings", trust_remote_code=True)
model = AutoModel.from_pretrained("neuml/bert-hash-nano-embeddings", trust_remote_code=True)
# Tokenize sentences
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')
# Compute token embeddingswith torch.no_grad():
output = model(**inputs)
# Perform pooling. In this case, mean pooling.
embeddings = meanpooling(output, inputs['attention_mask'])
print("Sentence embeddings:")
print(embeddings)
In analyzing the results,
bert-hash-nano-embeddings
is better across the board vs MUVERA with
colbert-muvera-nano
. It keeps
98%
of the performance of full multi-vector maxsim vs
95%
for MUVERA. Comparing the standard MUVERA output of
10240
vs
128
dimensions,
10K
standard F32 vectors needs
400 MB
of storage vs
5 MB
By itself for a
970K
parameter model, the scores are really good. When paired with re-ranking with a
970K
ColBERT model, the scores are even better. Competitive with common small models as shown above at only
~4%
of the number of parameters.
While this isn't a state of the art model, it's an extremely competitive method for building vectors on edge and low resource devices.
bert-hash-nano-embeddings huggingface.co is an AI model on huggingface.co that provides bert-hash-nano-embeddings's model effect (), which can be used instantly with this NeuML bert-hash-nano-embeddings model. huggingface.co supports a free trial of the bert-hash-nano-embeddings model, and also provides paid use of the bert-hash-nano-embeddings. Support call bert-hash-nano-embeddings model through api, including Node.js, Python, http.
bert-hash-nano-embeddings huggingface.co is an online trial and call api platform, which integrates bert-hash-nano-embeddings's modeling effects, including api services, and provides a free online trial of bert-hash-nano-embeddings, you can try bert-hash-nano-embeddings online for free by clicking the link below.
NeuML bert-hash-nano-embeddings online free url in huggingface.co:
bert-hash-nano-embeddings is an open source model from GitHub that offers a free installation service, and any user can find bert-hash-nano-embeddings on GitHub to install. At the same time, huggingface.co provides the effect of bert-hash-nano-embeddings install, users can directly use bert-hash-nano-embeddings installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
bert-hash-nano-embeddings install url in huggingface.co: