This is a
H.G. BERT Small
model fined-tuned using
sentence-transformers
. It maps sentences & paragraphs to a 384 dimensional dense vector space and can be used for tasks like clustering or semantic search.
The training dataset was generating using a random sample of
Historical English Books
title-abstract pairs and query-sentences pairs.
This model was trained using the following three step process.
Train a new model using the generated dataset with hard negatives
Unlike other models in this small domain model series, this model did not use distillation. This prevents against data leakage and learning modern similarity patterns. This keeps the model focused on 1899 and prior.
Usage (txtai)
This model can be used to build embeddings databases with
txtai
for semantic search and/or as a knowledge source for retrieval augmented generation (RAG).
import txtai
embeddings = txtai.Embeddings(path="neuml/hgbert-small-embeddings", content=True)
embeddings.index(documents())
# Run a query
embeddings.search("query to run")
from sentence_transformers import SentenceTransformer
sentences = ["This is an example sentence", "Each sentence is converted"]
model = SentenceTransformer("neuml/hgbert-small-embeddings")
embeddings = model.encode(sentences)
print(embeddings)
Usage (Hugging Face Transformers)
The model can also be used directly with Transformers.
from transformers import AutoTokenizer, AutoModel
import torch
# Mean Pooling - Take attention mask into account for correct averagingdefmeanpooling(output, mask):
embeddings = output[0] # First element of model_output contains all token embeddings
mask = mask.unsqueeze(-1).expand(embeddings.size()).float()
return torch.sum(embeddings * mask, 1) / torch.clamp(mask.sum(1), min=1e-9)
# Sentences we want sentence embeddings for
sentences = ['This is an example sentence', 'Each sentence is converted']
# Load model from HuggingFace Hub
tokenizer = AutoTokenizer.from_pretrained("neuml/hgbert-small-embeddings")
model = AutoModel.from_pretrained("neuml/hgbert-small-embeddings")
# Tokenize sentences
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')
# Compute token embeddingswith torch.no_grad():
output = model(**inputs)
# Perform pooling. In this case, mean pooling.
embeddings = meanpooling(output, inputs['attention_mask'])
print("Sentence embeddings:")
print(embeddings)
This model is a solid performer at a small size. It beats all models by a comfortable margin including ones with 8 billion parameters!
It can be used in CPU-only setups without trading off much on the accuracy front. It shows how small models can excel at specialized domains, requiring less compute and disk space.
hgbert-small-embeddings huggingface.co is an AI model on huggingface.co that provides hgbert-small-embeddings's model effect (), which can be used instantly with this NeuML hgbert-small-embeddings model. huggingface.co supports a free trial of the hgbert-small-embeddings model, and also provides paid use of the hgbert-small-embeddings. Support call hgbert-small-embeddings model through api, including Node.js, Python, http.
hgbert-small-embeddings huggingface.co is an online trial and call api platform, which integrates hgbert-small-embeddings's modeling effects, including api services, and provides a free online trial of hgbert-small-embeddings, you can try hgbert-small-embeddings online for free by clicking the link below.
NeuML hgbert-small-embeddings online free url in huggingface.co:
hgbert-small-embeddings is an open source model from GitHub that offers a free installation service, and any user can find hgbert-small-embeddings on GitHub to install. At the same time, huggingface.co provides the effect of hgbert-small-embeddings install, users can directly use hgbert-small-embeddings installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
hgbert-small-embeddings install url in huggingface.co: