The Latest AIs, every day
AIs with the most favorites on Toolify
AIs with the highest website traffic (monthly visits)
AI Tools by Apps
Discover the Discord of AI
AI Tools by browser extensions
GPTs from GPT Store
Discover The Best Model For AI
Top AI lists by month and monthly visits.
Top AI lists by category and monthly visits.
Top AI lists by region and monthly visits.
Top AI lists by source and monthly visits.
Top AI lists by revenue and real traffic.

elephant-embeddings-v1-multimodal-large
is the large multimodal embedding model in the
Agentic Intelligence Lab Elephant Embeddings V1
family.
This ModelScope release is maintained by
agentic-intelligence-lab
to make Elephant embedding models easier to download and deploy in mainland China. It mirrors and renames the upstream HuggingFace model
llm-semantic-router/multi-modal-embed-large
under a consistent Elephant model namespace.
This model is a production-oriented multimodal embedding model for semantic routing, retrieval, and cross-modal matching across text, image, and audio.
It is not a generative chat or captioning model. Instead, it maps different modalities into one shared embedding space so agent systems can compare requests, screenshots, documents, and audio records with the same retrieval interface.
| Item | Value |
|---|---|
| Family | Elephant Embeddings V1 |
| Maintainer | Agentic Intelligence Lab |
| Model type | Multimodal embedding model |
| Modalities | Text, image, audio |
| Architecture | Custom PyTorch tri-encoder |
| Text encoder |
llm-semantic-router/mmbert-embed-32k-2d-matryoshka
|
| Image encoder |
google/siglip2-so400m-patch14-384
|
| Audio encoder |
openai/whisper-medium
|
| Embedding dimension | 768 |
| Max text length | 32,768 tokens |
| Objective | Cached multiple negatives ranking loss |
| Upstream source |
llm-semantic-router/multi-modal-embed-large
|
| License | Apache 2.0 |
Agentic products increasingly need to retrieve and route over mixed inputs: user text, screenshots, UI states, documents, voice notes, support calls, and multimodal memory. This model is designed for that operating pattern.
Key advantages:
| Scenario | Example |
|---|---|
| Multimodal RAG | Retrieve text notes using an image or audio query |
| Agent routing | Route screenshots, user text, or voice requests to the right tool or workflow |
| Memory search | Search mixed text/image/audio memory stores in one vector space |
| Support and operations | Match tickets, screenshots, logs, and recorded calls semantically |
| Offline indexing | Build high-quality 768d multimodal indexes |
pip install modelscope torch sentence-transformers transformers accelerate safetensors pillow librosa soundfile
import json
import os
import sys
import torch
import torch.nn.functional as F
from modelscope import snapshot_download
repo_id = "agentic-intelligence-lab/elephant-embeddings-v1-multimodal-large"
local_dir = snapshot_download(repo_id)
sys.path.insert(0, os.path.join(local_dir, "src"))
from hf_st_mm.data import PairItem
from hf_st_mm.model import MultiModalSentenceEmbedder
with open(os.path.join(local_dir, "config.json"), "r", encoding="utf-8") as handle:
cfg = json.load(handle)
model = MultiModalSentenceEmbedder(
text_encoder_name=cfg["model"]["text_encoder_name"],
image_encoder_name=cfg["model"]["image_encoder_name"],
audio_encoder_name=cfg["model"]["audio_encoder_name"],
embedding_dim=int(cfg["model"]["embedding_dim"]),
max_text_length=int(cfg["model"]["max_text_length"]),
)
state_dict = torch.load(os.path.join(local_dir, "model.pt"), map_location="cpu")
model.load_state_dict(state_dict)
model.eval()
items = [
PairItem(modality="text", value="route this request to the billing workflow"),
PairItem(modality="image", value="/path/to/screenshot.png"),
PairItem(modality="audio", value="/path/to/call.wav"),
]
with torch.no_grad():
embeddings = model.encode_items(items)
print(embeddings.shape) # [3, 768]
query = PairItem(modality="text", value="refund request for a wrong charge")
candidate = PairItem(modality="audio", value="/path/to/refund_call.wav")
with torch.no_grad():
embs = model.encode_items([query, candidate])
similarity = F.cosine_similarity(embs[0:1], embs[1:2]).item()
print(f"similarity={similarity:.4f}")
| Metric | Value |
|---|---|
| Eval loss | 0.389702 |
| Eval top1 | 0.861707 |
The validation metrics come from the tri-encoder cached retrieval validation path used during export. They are intended as a release sanity snapshot rather than a public leaderboard claim.
| File | Description |
|---|---|
model.pt
|
Exported PyTorch weights |
config.json
|
Tri-encoder and training/export configuration |
src/hf_st_mm/
|
Python package used to construct and run the model |
README.md
|
This model card |
This ModelScope package is published by
agentic-intelligence-lab
as part of the Elephant model release line. It mirrors the upstream HuggingFace model
llm-semantic-router/multi-modal-embed-large
and keeps the model artifacts unchanged except for the repository naming and model card presentation.
hf_st_mm
source code.
@misc{elephant-embeddings-v1-multimodal-large,
title={Elephant Embeddings V1 Multimodal Large},
author={Agentic Intelligence Lab},
year={2026},
url={https://modelscope.cn/models/agentic-intelligence-lab/elephant-embeddings-v1-multimodal-large}
}
Apache 2.0
elephant-embeddings-v1-multimodal-large huggingface.co is an AI model on huggingface.co that provides elephant-embeddings-v1-multimodal-large's model effect (), which can be used instantly with this agentic-in elephant-embeddings-v1-multimodal-large model. huggingface.co supports a free trial of the elephant-embeddings-v1-multimodal-large model, and also provides paid use of the elephant-embeddings-v1-multimodal-large. Support call elephant-embeddings-v1-multimodal-large model through api, including Node.js, Python, http.
elephant-embeddings-v1-multimodal-large huggingface.co is an online trial and call api platform, which integrates elephant-embeddings-v1-multimodal-large's modeling effects, including api services, and provides a free online trial of elephant-embeddings-v1-multimodal-large, you can try elephant-embeddings-v1-multimodal-large online for free by clicking the link below.
elephant-embeddings-v1-multimodal-large is an open source model from GitHub that offers a free installation service, and any user can find elephant-embeddings-v1-multimodal-large on GitHub to install. At the same time, huggingface.co provides the effect of elephant-embeddings-v1-multimodal-large install, users can directly use elephant-embeddings-v1-multimodal-large installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
