2.3B parameters.
Text, speech, general audio, image, video, and visually-rich documents
in a single shared cosine space, from a backbone that is
never updated
.
Omni-Embed-Mini recasts cross-modal alignment as
self-distillation through one shared frozen
causal backbone
. Every media sample is paired with a dense cascaded caption; the teacher target
is the EOS-pooled embedding of that caption produced by the
identical
frozen backbone that
processes the student input. Teacher and student therefore inhabit byte-identical geometry, and
because no text-side parameter is ever updated, adding audio cannot degrade the inherited
text and vision representations.
Only the audio projectors and small phased LoRA adapters on the audio encoders are trained. A
Matryoshka SigLIP contrastive objective and an online hybrid hard-negative miner supply the
contrastive signal. This variant keeps the Qwen3-VL
native
visual tower, so image, video, and
document-page quality is inherited intact from the backbone while speech and audio are added on top.
Results
Six modalities, evaluated with the pipeline in the
code repository
.
Modality
Benchmark
Metric
2.3B v1
2.3B v2
0.9B v1
0.9B v2
Text
MTEB-v2 BEIR-8
nDCG@10
47.94
49.57
Speech
MAEB (12 tasks)
mean
48.86
43.28
Audio
MAEB (10 tasks)
mean
33.44
33.43
Image
MMEB-V2 (10 tasks)
hit@1
64.80
26.29
Video
MMEB-V2 (6 tasks)
hit@1
55.18
18.48
Vis-Doc
ViDoRe-v3 (7 tasks)
nDCG@5
58.10
46.92
v1
and
v2
are tagged revisions of this same repository, so
revision="v1.0"
pins the
numbers in the v1 column.
import torch
from transformers import AutoModel, AutoProcessor
REPO = "MBZUAI/Omni-Embed-Mini-2.3B"
model = AutoModel.from_pretrained(
REPO, revision="v1.0", trust_remote_code=True, dtype=torch.bfloat16,
).cuda().eval()
processor = AutoProcessor.from_pretrained(REPO, revision="v1.0", trust_remote_code=True)
# OpenAI-style multimodal messages: text, audio, image, video, doc, or any composition.
messages = [{"role": "user", "content": [
{"type": "audio", "audio": "path/to/clip.wav"},
{"type": "text", "text": "rain on a tin roof at night"},
]}]
inputs = processor.apply_chat_template(
messages, role="passage", tokenize=True, return_tensors="pt",
).to("cuda")
with torch.no_grad():
doc = model(**inputs).pooler_output # (1, 2048), already L2-normalised# Query side. role="query" applies the retrieval instruction template.# text_recipe="chat" puts a text query in the same subspace as media documents;# omit it only for pure text-to-text retrieval, which uses the native recipe.
query = processor.apply_chat_template(
[{"role": "user", "content": "rain on a tin roof at night"}],
role="query", tokenize=True, return_tensors="pt", text_recipe="chat",
).to("cuda")
with torch.no_grad():
q = model(**query).pooler_output
print("cosine:", float(q @ doc.T)) # normalised, so dot == cosine
Matryoshka (truncatable) embeddings
model(**inputs, truncate_dim=N).pooler_output
returns a prefix-truncated, renormalised
vector. Slicing
pooler_output
yourself is not equivalent unless you renormalise after.
Supported
N
: 128, 256, 512, 1024, 2048. Use a smaller
N
to cut index size at a modest recall cost.
Architecture
Component
This model
Backbone (frozen)
Qwen/Qwen3-VL-Embedding-2B
Embedding dim
2048
Vision
native
Qwen3-VL visual tower (inside
backbone/
; no separate file)
Speech encoder
openai/whisper-small
Audio encoder
mispeech/dasheng-base
Matryoshka dims
128, 256, 512, 1024, 2048
Video
native video path (
video_as_images: false
)
Trained parameters: audio projectors plus phase-2 LoRA adapters on the audio encoders
(48 Whisper and 24 Dasheng adapted tensors, merged into the released weights). The backbone is
frozen at every stage and ships unmodified apart from a vocabulary resize that adds the media
placeholder tokens.
There is no
vision_encoder.pt
, which is expected for this variant,
whose vision weights live inside
backbone/
.
Omni-Embed-Mini-2.3B huggingface.co is an AI model on huggingface.co that provides Omni-Embed-Mini-2.3B's model effect (), which can be used instantly with this MBZUAI Omni-Embed-Mini-2.3B model. huggingface.co supports a free trial of the Omni-Embed-Mini-2.3B model, and also provides paid use of the Omni-Embed-Mini-2.3B. Support call Omni-Embed-Mini-2.3B model through api, including Node.js, Python, http.
Omni-Embed-Mini-2.3B huggingface.co is an online trial and call api platform, which integrates Omni-Embed-Mini-2.3B's modeling effects, including api services, and provides a free online trial of Omni-Embed-Mini-2.3B, you can try Omni-Embed-Mini-2.3B online for free by clicking the link below.
MBZUAI Omni-Embed-Mini-2.3B online free url in huggingface.co:
Omni-Embed-Mini-2.3B is an open source model from GitHub that offers a free installation service, and any user can find Omni-Embed-Mini-2.3B on GitHub to install. At the same time, huggingface.co provides the effect of Omni-Embed-Mini-2.3B install, users can directly use Omni-Embed-Mini-2.3B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Omni-Embed-Mini-2.3B install url in huggingface.co: