This is an
AstroBERT Small
model fined-tuned using
sentence-transformers
. It maps sentences & paragraphs to a 384 dimensional dense vector space and can be used for tasks like clustering or semantic search.
The training dataset was generated using a random sample of
ArXiv abstracts
labeled as
astro-ph
.
This model can be used to build embeddings databases with
txtai
for semantic search and/or as a knowledge source for retrieval augmented generation (RAG).
import txtai
embeddings = txtai.Embeddings(path="neuml/astrobert-small-embeddings", content=True)
embeddings.index(documents())
# Run a query
embeddings.search("query to run")
from sentence_transformers import SentenceTransformer
sentences = ["This is an example sentence", "Each sentence is converted"]
model = SentenceTransformer("neuml/astrobert-small-embeddings")
embeddings = model.encode(sentences)
print(embeddings)
Usage (Hugging Face Transformers)
The model can also be used directly with Transformers.
from transformers import AutoTokenizer, AutoModel
import torch
# Mean Pooling - Take attention mask into account for correct averagingdefmeanpooling(output, mask):
embeddings = output[0] # First element of model_output contains all token embeddings
mask = mask.unsqueeze(-1).expand(embeddings.size()).float()
return torch.sum(embeddings * mask, 1) / torch.clamp(mask.sum(1), min=1e-9)
# Sentences we want sentence embeddings for
sentences = ['This is an example sentence', 'Each sentence is converted']
# Load model from HuggingFace Hub
tokenizer = AutoTokenizer.from_pretrained("neuml/astrobert-small-embeddings")
model = AutoModel.from_pretrained("neuml/astrobert-small-embeddings")
# Tokenize sentences
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')
# Compute token embeddingswith torch.no_grad():
output = model(**inputs)
# Perform pooling. In this case, mean pooling.
embeddings = meanpooling(output, inputs['attention_mask'])
print("Sentence embeddings:")
print(embeddings)
Evaluation Results
A
BEIR-compatible dataset
was generated to facilitate the evaluation process. This is a separate random sample of Wikipedia articles alongside generated user queries.
Evaluation results are shown below.
NDCG
is used as the evaluation metric.
This model is a solid performer at a small size. It beats the same sized
all-MiniLM-L6-v2
model by a significant margin. It beats the 600M parameter Qwen3 Embeddings model which is over 25x larger. It scores slightly lower than the model it's distilled from (
Qwen3-Embedding-8B
).
This is a great model that can be used in CPU-only setups without trading off much on the accuracy front. It shows how small models can excel at specialized domains, requiring less compute and disk space.
astrobert-small-embeddings huggingface.co is an AI model on huggingface.co that provides astrobert-small-embeddings's model effect (), which can be used instantly with this NeuML astrobert-small-embeddings model. huggingface.co supports a free trial of the astrobert-small-embeddings model, and also provides paid use of the astrobert-small-embeddings. Support call astrobert-small-embeddings model through api, including Node.js, Python, http.
astrobert-small-embeddings huggingface.co is an online trial and call api platform, which integrates astrobert-small-embeddings's modeling effects, including api services, and provides a free online trial of astrobert-small-embeddings, you can try astrobert-small-embeddings online for free by clicking the link below.
NeuML astrobert-small-embeddings online free url in huggingface.co:
astrobert-small-embeddings is an open source model from GitHub that offers a free installation service, and any user can find astrobert-small-embeddings on GitHub to install. At the same time, huggingface.co provides the effect of astrobert-small-embeddings install, users can directly use astrobert-small-embeddings installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
astrobert-small-embeddings install url in huggingface.co: