nomic-embed-text-v1.5: Resizable Production Embeddings with Matryoshka Representation Learning
nomic-embed-text-v1.5
is an improvement upon
Nomic Embed
that utilizes
Matryoshka Representation Learning
which gives developers the flexibility to trade off the embedding size for a negligible reduction in performance.
Name
SeqLen
Dimension
MTEB
nomic-embed-text-v1
8192
768
62.39
nomic-embed-text-v1.5
8192
768
62.28
nomic-embed-text-v1.5
8192
512
61.96
nomic-embed-text-v1.5
8192
256
61.04
nomic-embed-text-v1.5
8192
128
59.34
nomic-embed-text-v1.5
8192
64
56.10
Exciting Update!
:
nomic-embed-text-v1.5
is now multimodal!
nomic-embed-vision-v1
is aligned to the embedding space of
nomic-embed-text-v1.5
, meaning any text embedding is multimodal!
Hosted Inference API
The easiest way to get started with Nomic Embed is through the Nomic Embedding API.
Generating embeddings with the
nomic
Python client is as easy as
Click the Nomic Atlas map below to visualize a 5M sample of our contrastive pretraining data!
Training Details
We train our embedder using a multi-stage training pipeline. Starting from a long-context
BERT model
,
the first unsupervised contrastive stage trains on a dataset generated from weakly related text pairs, such as question-answer pairs from forums like StackExchange and Quora, title-body pairs from Amazon reviews, and summarizations from news articles.
In the second finetuning stage, higher quality labeled datasets such as search queries and answers from web searches are leveraged. Data curation and hard-example mining is crucial in this stage.
Training data to train the models is released in its entirety. For more details, see the
contrastors
repository
Usage
Note
nomic-embed-text
requires
prefixes! We support the prefixes
[search_query, search_document, classification, clustering]
.
For retrieval applications, you should prepend
search_document
for all your documents and
search_query
for your queries.
For example, you are building a RAG application over the top of Wikipedia. You would embed all Wikipedia articles with the prefix
search_document
and any questions you ask with
search_query
. For example:
queries = ["search_query: who is the first president of the united states?", "search_query: when was babe ruth born?"]
documents = ["search_document: <article about US Presidents>", "search_document: <article about Babe Ruth>"]
Sentence Transformers
import torch.nn.functional as F
from sentence_transformers import SentenceTransformer
matryoshka_dim = 512
model = SentenceTransformer("nomic-ai/nomic-embed-text-v1.5", trust_remote_code=True)
sentences = ['search_query: What is TSNE?', 'search_query: Who is Laurens van der Maaten?']
embeddings = model.encode(sentences, convert_to_tensor=True)
embeddings = F.layer_norm(embeddings, normalized_shape=(embeddings.shape[1],))
embeddings = embeddings[:, :matryoshka_dim]
embeddings = F.normalize(embeddings, p=2, dim=1)
print(embeddings)
Transformers
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel
def mean_pooling(model_output, attention_mask):
token_embeddings = model_output[0]
input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(input_mask_expanded.sum(1), min=1e-9)
sentences = ['search_query: What is TSNE?', 'search_query: Who is Laurens van der Maaten?']
tokenizer = AutoTokenizer.from_pretrained('bert-base-uncased')
model = AutoModel.from_pretrained('nomic-ai/nomic-embed-text-v1.5', trust_remote_code=True, safe_serialization=True)
model.eval()
encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')
+ matryoshka_dim = 512
with torch.no_grad():
model_output = model(**encoded_input)
embeddings = mean_pooling(model_output, encoded_input['attention_mask'])
+ embeddings = F.layer_norm(embeddings, normalized_shape=(embeddings.shape[1],))+ embeddings = embeddings[:, :matryoshka_dim]
embeddings = F.normalize(embeddings, p=2, dim=1)
print(embeddings)
The model natively supports scaling of the sequence length past 2048 tokens. To do so,
- tokenizer = AutoTokenizer.from_pretrained('bert-base-uncased')+ tokenizer = AutoTokenizer.from_pretrained('bert-base-uncased', model_max_length=8192)- model = AutoModel.from_pretrained('nomic-ai/nomic-embed-text-v1', trust_remote_code=True)+ model = AutoModel.from_pretrained('nomic-ai/nomic-embed-text-v1', trust_remote_code=True, rotary_scaling_factor=2)
Transformers.js
import { pipeline, layer_norm } from'@xenova/transformers';
// Create a feature extraction pipelineconst extractor = awaitpipeline('feature-extraction', 'nomic-ai/nomic-embed-text-v1.5', {
quantized: false, // Comment out this line to use the quantized version
});
// Define sentencesconst texts = ['search_query: What is TSNE?', 'search_query: Who is Laurens van der Maaten?'];
// Compute sentence embeddingslet embeddings = awaitextractor(texts, { pooling: 'mean' });
console.log(embeddings); // Tensor of shape [2, 768]const matryoshka_dim = 512;
embeddings = layer_norm(embeddings, [embeddings.dims[1]])
.slice(null, [0, matryoshka_dim])
.normalize(2, -1);
console.log(embeddings.tolist());
If you find the model, dataset, or training code useful, please cite our work
@misc{nussbaum2024nomic,
title={Nomic Embed: Training a Reproducible Long Context Text Embedder},
author={Zach Nussbaum and John X. Morris and Brandon Duderstadt and Andriy Mulyar},
year={2024},
eprint={2402.01613},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
Runs of nomic-ai nomic-embed-text-v1.5 on huggingface.co
15.9M
Total runs
0
24-hour runs
-146.7K
3-day runs
-303.0K
7-day runs
700.8K
30-day runs
More Information About nomic-embed-text-v1.5 huggingface.co Model
nomic-embed-text-v1.5 huggingface.co is an AI model on huggingface.co that provides nomic-embed-text-v1.5's model effect (), which can be used instantly with this nomic-ai nomic-embed-text-v1.5 model. huggingface.co supports a free trial of the nomic-embed-text-v1.5 model, and also provides paid use of the nomic-embed-text-v1.5. Support call nomic-embed-text-v1.5 model through api, including Node.js, Python, http.
nomic-embed-text-v1.5 huggingface.co is an online trial and call api platform, which integrates nomic-embed-text-v1.5's modeling effects, including api services, and provides a free online trial of nomic-embed-text-v1.5, you can try nomic-embed-text-v1.5 online for free by clicking the link below.
nomic-ai nomic-embed-text-v1.5 online free url in huggingface.co:
nomic-embed-text-v1.5 is an open source model from GitHub that offers a free installation service, and any user can find nomic-embed-text-v1.5 on GitHub to install. At the same time, huggingface.co provides the effect of nomic-embed-text-v1.5 install, users can directly use nomic-embed-text-v1.5 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
nomic-embed-text-v1.5 install url in huggingface.co: