The crispy sentence embedding family from
Mixedbread
.
🍞 Looking for a simple end-to-end retrieval solution? Meet Omni, our multimodal and multilingual model.
Get in touch for access.
mixedbread-ai/mxbai-embed-large-v1
Here, we provide several ways to produce sentence embeddings. Please note that you have to provide the prompt
Represent this sentence for searching relevant passages:
for query if you want to use it for retrieval. Besides that you don't need any prompt. Our model also supports
Matryoshka Representation Learning and binary quantization
.
Quickstart
Here, we provide several ways to produce sentence embeddings. Please note that you have to provide the prompt
Represent this sentence for searching relevant passages:
for query if you want to use it for retrieval. Besides that you don't need any prompt.
sentence-transformers
python -m pip install -U sentence-transformers
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
from sentence_transformers.quantization import quantize_embeddings
# 1. Specify preffered dimensions
dimensions = 512# 2. load model
model = SentenceTransformer("mixedbread-ai/mxbai-embed-large-v1", truncate_dim=dimensions)
# The prompt used for query retrieval tasks:# query_prompt = 'Represent this sentence for searching relevant passages: '
query = "A man is eating a piece of bread"
docs = [
"A man is eating food.",
"A man is eating pasta.",
"The girl is carrying a baby.",
"A man is riding a horse.",
]
# 2. Encode
query_embedding = model.encode(query, prompt_name="query")
# Equivalent Alternatives:# query_embedding = model.encode(query_prompt + query)# query_embedding = model.encode(query, prompt=query_prompt)
docs_embeddings = model.encode(docs)
# Optional: Quantize the embeddings
binary_query_embedding = quantize_embeddings(query_embedding, precision="ubinary")
binary_docs_embeddings = quantize_embeddings(docs_embeddings, precision="ubinary")
similarities = cos_sim(query_embedding, docs_embeddings)
print('similarities:', similarities)
Transformers
from typing importDictimport torch
import numpy as np
from transformers import AutoModel, AutoTokenizer
from sentence_transformers.util import cos_sim
# For retrieval you need to pass this prompt. Please find our more in our blog post.deftransform_query(query: str) -> str:
""" For retrieval, add the prompt for query (not for documents). """returnf'Represent this sentence for searching relevant passages: {query}'# The model works really well with cls pooling (default) but also with mean pooling.defpooling(outputs: torch.Tensor, inputs: Dict, strategy: str = 'cls') -> np.ndarray:
if strategy == 'cls':
outputs = outputs[:, 0]
elif strategy == 'mean':
outputs = torch.sum(
outputs * inputs["attention_mask"][:, :, None], dim=1) / torch.sum(inputs["attention_mask"], dim=1, keepdim=True)
else:
raise NotImplementedError
return outputs.detach().cpu().numpy()
# 1. load model
model_id = 'mixedbread-ai/mxbai-embed-large-v1'
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id).cuda()
docs = [
transform_query('A man is eating a piece of bread'),
"A man is eating food.",
"A man is eating pasta.",
"The girl is carrying a baby.",
"A man is riding a horse.",
]
# 2. encode
inputs = tokenizer(docs, padding=True, return_tensors='pt')
for k, v in inputs.items():
inputs[k] = v.cuda()
outputs = model(**inputs).last_hidden_state
embeddings = pooling(outputs, inputs, 'cls')
similarities = cos_sim(embeddings[0], embeddings[1:])
print('similarities:', similarities)
Transformers.js
If you haven't already, you can install the
Transformers.js
JavaScript library from
NPM
using:
npm i @xenova/transformers
You can then use the model to compute embeddings like this:
import { pipeline, cos_sim } from'@xenova/transformers';
// Create a feature extraction pipelineconst extractor = awaitpipeline('feature-extraction', 'mixedbread-ai/mxbai-embed-large-v1', {
quantized: false, // Comment out this line to use the quantized version
});
// Generate sentence embeddingsconst docs = [
'Represent this sentence for searching relevant passages: A man is eating a piece of bread',
'A man is eating food.',
'A man is eating pasta.',
'The girl is carrying a baby.',
'A man is riding a horse.',
]
const output = awaitextractor(docs, { pooling: 'cls' });
// Compute similarity scoresconst [source_embeddings, ...document_embeddings ] = output.tolist();
const similarities = document_embeddings.map(x =>cos_sim(source_embeddings, x));
console.log(similarities); // [0.7919578577247139, 0.6369278664248345, 0.16512018371357193, 0.3620778366720027]
Using API
You can use the model via our API as follows:
from mixedbread_ai.client import MixedbreadAI, EncodingFormat
from sklearn.metrics.pairwise import cosine_similarity
import os
mxbai = MixedbreadAI(api_key="{MIXEDBREAD_API_KEY}")
english_sentences = [
'What is the capital of Australia?',
'Canberra is the capital of Australia.'
]
res = mxbai.embeddings(
input=english_sentences,
model="mixedbread-ai/mxbai-embed-large-v1",
normalized=True,
encoding_format=[EncodingFormat.FLOAT, EncodingFormat.UBINARY, EncodingFormat.INT_8],
dimensions=512
)
encoded_embeddings = res.data[0].embedding
print(res.dimensions, encoded_embeddings.ubinary, encoded_embeddings.float_, encoded_embeddings.int_8)
The API comes with native int8 and binary quantization support! Check out the
docs
for more information.
As of March 2024, our model archives SOTA performance for Bert-large sized models on the
MTEB
. It ourperforms commercial models like OpenAIs text-embedding-3-large and matches the performance of model 20x it's size like the
echo-mistral-7b
. Our model was trained with no overlap of the MTEB data, which indicates that our model generalizes well across several domains, tasks and text length. We know there are some limitations with this model, which will be fixed in v2.
Embeddings in their commonly used form (float arrays) have a high memory footprint when used at scale. Two approaches to solve this problem are Matryoshka Representation Learning (MRL) and (Binary) Quantization. While MRL reduces the number of dimensions of an embedding, binary quantization transforms the value of each dimension from a float32 into a lower precision (int8 or even binary).
The model supports both approaches!
You can also take it one step further, and combine both MRL and quantization. This combination of binary quantization and MRL allows you to reduce the memory usage of your embeddings significantly. This leads to much lower costs when using a vector database in particular. You can read more about the technology and its advantages in our
blog post
.
Community
Please join our
Discord Community
and share your feedback and thoughts! We are here to help and also always happy to chat.
License
Apache 2.0
Citation
@online{emb2024mxbai,
title={Open Source Strikes Bread - New Fluffy Embeddings Model},
author={Sean Lee and Aamir Shakir and Darius Koenig and Julius Lipp},
year={2024},
url={https://www.mixedbread.ai/blog/mxbai-embed-large-v1},
}
@article{li2023angle,
title={AnglE-optimized Text Embeddings},
author={Li, Xianming and Li, Jing},
journal={arXiv preprint arXiv:2309.12871},
year={2023}
}
Runs of OrcaDB mxbai-large on huggingface.co
847
Total runs
0
24-hour runs
-41
3-day runs
-20
7-day runs
-363
30-day runs
More Information About mxbai-large huggingface.co Model
mxbai-large huggingface.co is an AI model on huggingface.co that provides mxbai-large's model effect (), which can be used instantly with this OrcaDB mxbai-large model. huggingface.co supports a free trial of the mxbai-large model, and also provides paid use of the mxbai-large. Support call mxbai-large model through api, including Node.js, Python, http.
mxbai-large huggingface.co is an online trial and call api platform, which integrates mxbai-large's modeling effects, including api services, and provides a free online trial of mxbai-large, you can try mxbai-large online for free by clicking the link below.
OrcaDB mxbai-large online free url in huggingface.co:
mxbai-large is an open source model from GitHub that offers a free installation service, and any user can find mxbai-large on GitHub to install. At the same time, huggingface.co provides the effect of mxbai-large install, users can directly use mxbai-large installed effect in huggingface.co for debugging and trial. It also supports api for free installation.