Youtu-Embedding
is a state-of-the-art, general-purpose text embedding model developed by Tencent Youtu Lab. It delivers exceptional performance across a wide range of natural language processing tasks, including Information Retrieval (IR), Semantic Textual Similarity (STS), Clustering, Reranking, and Classification.
Top-Ranked Performance
: Achieved the #1 score of
77.46
on the authoritative CMTEB (Chinese Massive Text Embedding Benchmark) as of September 2025, demonstrating its powerful and robust text representation capabilities.
Innovative Training Framework
: Features a Collaborative-Discriminative Fine-tuning Framework designed to resolve the "negative transfer" problem in multi-task learning. This is accomplished through a unified data format, task-differentiated loss functions, and a dynamic single-task sampling mechanism.
Note
: You can easily adapt and fine-tune the model on your own datasets for domain-specific tasks. For implementation details, please refer to the
training code
.
import torch
from langchain.docstore.document import Document
from langchain_community.vectorstores import FAISS
from langchain_huggingface.embeddings import HuggingFaceEmbeddings
model_name_or_path = "tencent/Youtu-Embedding"
device = "cuda"if torch.cuda.is_available() else"cpu"
model_kwargs = {
'trust_remote_code': True,
'device': device
}
embedder = HuggingFaceEmbeddings(
model_name=model_name_or_path,
model_kwargs=model_kwargs,
)
query_instruction = "Instruction: Given a search query, retrieve passages that answer the question \nQuery: "
doc_instruction = ""
data = [
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
"Saturn, famous for its rings, is sometimes mistaken for the Red Planet."
]
documents = [Document(page_content=text, metadata={"id": i}) for i, text inenumerate(data)]
vector_store = FAISS.from_documents(documents, embedder, distance_strategy="MAX_INNER_PRODUCT")
query = "Which planet is known as the Red Planet?"
instructed_query = query_instruction + query
results = vector_store.similarity_search_with_score(instructed_query, k=3)
print(f"Original Query: {query}\n")
print("Results:")
for doc, score in results:
print(f"- Text: {doc.page_content} (Score: {score:.4f})")
4. Using
LlamaIndex
🦙
This is perfect for integrating the model into your
LlamaIndex
search and retrieval systems.
import faiss
import torch
from llama_index.core.schema import TextNode
from llama_index.core.vector_stores import VectorStoreQuery
from llama_index.vector_stores.faiss import FaissVectorStore
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
model_name_or_path = "tencent/Youtu-Embedding"
device = "cuda"if torch.cuda.is_available() else"cpu"
embeddings = HuggingFaceEmbedding(
model_name=model_name_or_path,
trust_remote_code=True,
device=device,
query_instruction="Instruction: Given a search query, retrieve passages that answer the question \nQuery: ",
text_instruction=""
)
data = [
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
"Saturn, famous for its rings, is sometimes mistaken for the Red Planet."
]
nodes = [TextNode(id_=str(i), text=text) for i, text inenumerate(data)]
for node in nodes:
node.embedding = embeddings.get_text_embedding(node.get_content())
embed_dim = len(nodes[0].embedding)
store = FaissVectorStore(faiss_index=faiss.IndexFlatIP(embed_dim))
store.add(nodes)
query = "Which planet is known as the Red Planet?"
query_embedding = embeddings.get_query_embedding(query)
results = store.query(
VectorStoreQuery(query_embedding=query_embedding, similarity_top_k=3)
)
print(f"Query: {query}\n")
print("Results:")
for idx, score inzip(results.ids, results.similarities):
print(f"- Text: {data[int(idx)]} (Score: {score:.4f})")
📊 CMTEB
Model
Param.
Mean(Task)
Mean(Type)
Class.
Clust.
Pair Class.
Rerank.
Retr.
STS
bge-multilingual-gemma2
9B
67.64
68.52
75.31
59.30
79.30
68.28
73.73
55.19
ritrieve_zh_v1
326M
72.71
73.85
76.88
66.50
85.98
72.86
76.97
63.92
Qwen3-Embedding-4B
4B
72.27
73.51
75.46
77.89
83.34
66.05
77.03
61.26
Qwen3-Embedding-8B
8B
73.84
75.00
76.97
80.08
84.23
66.99
78.21
63.53
Conan-embedding-v2
1.4B
74.24
75.99
76.47
68.84
92.44
74.41
78.31
65.48
Seed1.6-embedding
-
75.63
76.68
77.98
73.11
88.71
71.65
79.69
68.94
QZhou-Embedding
7B
76.99
78.58
79.99
70.91
95.07
74.85
78.80
71.89
Youtu-Embedding-V1
2B
77.60
78.85
78.04
79.67
89.69
73.85
80.95
70.91
Note
: Comparative scores are from the MTEB
leaderboard
, recorded on September 28, 2025.
🎉 Citation
@misc{zhang2025codiemb,
title={CoDiEmb: A Collaborative yet Distinct Framework for Unified Representation Learning in Information Retrieval and Semantic Textual Similarity},
author={Zhang, Bowen and Song, Zixin and Chen, Chunquan and Zhang, Qian-Wen and Yin, Di and Sun, Xing},
year={2025},
eprint={2508.11442},
archivePrefix={arXiv},
url={https://arxiv.org/abs/2508.11442},
}
Runs of tencent Youtu-Embedding on huggingface.co
552
Total runs
0
24-hour runs
31
3-day runs
-62
7-day runs
-1.7K
30-day runs
More Information About Youtu-Embedding huggingface.co Model
Youtu-Embedding huggingface.co is an AI model on huggingface.co that provides Youtu-Embedding's model effect (), which can be used instantly with this tencent Youtu-Embedding model. huggingface.co supports a free trial of the Youtu-Embedding model, and also provides paid use of the Youtu-Embedding. Support call Youtu-Embedding model through api, including Node.js, Python, http.
Youtu-Embedding huggingface.co is an online trial and call api platform, which integrates Youtu-Embedding's modeling effects, including api services, and provides a free online trial of Youtu-Embedding, you can try Youtu-Embedding online for free by clicking the link below.
tencent Youtu-Embedding online free url in huggingface.co:
Youtu-Embedding is an open source model from GitHub that offers a free installation service, and any user can find Youtu-Embedding on GitHub to install. At the same time, huggingface.co provides the effect of Youtu-Embedding install, users can directly use Youtu-Embedding installed effect in huggingface.co for debugging and trial. It also supports api for free installation.