BAAI / bge-multilingual-gemma2

huggingface.co
Total runs: 307.7K
24-hour runs: 4.6K
7-day runs: 23.3K
30-day runs: -25.5K
Model's Last Updated: October 13 2025
feature-extraction

Introduction of bge-multilingual-gemma2

Model Details of bge-multilingual-gemma2

FlagEmbedding

For more details please refer to our Github: FlagEmbedding .

BGE-Multilingual-Gemma2 is a LLM-based multilingual embedding model. It is trained on a diverse range of languages and tasks based on google/gemma-2-9b . BGE-Multilingual-Gemma2 primarily demonstrates the following advancements:

  • Diverse training data: The model's training data spans a broad range of languages, including English, Chinese, Japanese, Korean, French, and more.Additionally, the data covers a variety of task types, such as retrieval, classification, and clustering.
  • Outstanding performance: The model exhibits state-of-the-art (SOTA) results on multilingual benchmarks like MIRACL, MTEB-pl, and MTEB-fr. It also achieves excellent performance on other major evaluations, including MTEB, C-MTEB and AIR-Bench.
📑 Open-source Plan
  • Checkpoint
  • Training Data

We will release the training data of BGE-Multilingual-Gemma2 in the future.

Usage
Using FlagEmbedding
git clone https://github.com/FlagOpen/FlagEmbedding.git
cd FlagEmbedding
pip install -e .
from FlagEmbedding import FlagLLMModel
queries = ["how much protein should a female eat", "summit define"]
documents = [
    "As a general guideline, the CDC's average requirement of protein for women ages 19 to 70 is 46 grams per day. But, as you can see from this chart, you'll need to increase that if you're expecting or training for a marathon. Check out the chart below to see how much protein you should be eating each day.",
    "Definition of summit for English Language Learners. : 1  the highest point of a mountain : the top of a mountain. : 2  the highest level. : 3  a meeting or series of meetings between the leaders of two or more governments."
]
model = FlagLLMModel('BAAI/bge-multilingual-gemma2', 
                     query_instruction_for_retrieval="Given a web search query, retrieve relevant passages that answer the query.",
                     use_fp16=True) # Setting use_fp16 to True speeds up computation with a slight performance degradation
embeddings_1 = model.encode_queries(queries)
embeddings_2 = model.encode_corpus(documents)
similarity = embeddings_1 @ embeddings_2.T
print(similarity)
# [[ 0.559     0.01654 ]
# [-0.002575  0.4998  ]]

By default, FlagLLMModel will use all available GPUs when encoding. Please set os.environ["CUDA_VISIBLE_DEVICES"] to select specific GPUs. You also can set os.environ["CUDA_VISIBLE_DEVICES"]="" to make all GPUs unavailable.

Using Sentence Transformers
from sentence_transformers import SentenceTransformer
import torch

# Load the model, optionally in float16 precision for faster inference
model = SentenceTransformer("BAAI/bge-multilingual-gemma2", model_kwargs={"torch_dtype": torch.float16})

# Prepare a prompt given an instruction
instruction = 'Given a web search query, retrieve relevant passages that answer the query.'
prompt = f'<instruct>{instruction}\n<query>'
# Prepare queries and documents
queries = [
    'how much protein should a female eat',
    'summit define',
]
documents = [
    "As a general guideline, the CDC's average requirement of protein for women ages 19 to 70 is 46 grams per day. But, as you can see from this chart, you'll need to increase that if you're expecting or training for a marathon. Check out the chart below to see how much protein you should be eating each day.",
    "Definition of summit for English Language Learners. : 1  the highest point of a mountain : the top of a mountain. : 2  the highest level. : 3  a meeting or series of meetings between the leaders of two or more governments."
]

# Compute the query and document embeddings
query_embeddings = model.encode(queries, prompt=prompt)
document_embeddings = model.encode(documents)

# Compute the cosine similarity between the query and document embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[ 0.5591,  0.0164],
#         [-0.0026,  0.4993]], dtype=torch.float16)
Using HuggingFace Transformers
import torch
import torch.nn.functional as F

from torch import Tensor
from transformers import AutoTokenizer, AutoModel


def last_token_pool(last_hidden_states: Tensor,
                 attention_mask: Tensor) -> Tensor:
    left_padding = (attention_mask[:, -1].sum() == attention_mask.shape[0])
    if left_padding:
        return last_hidden_states[:, -1]
    else:
        sequence_lengths = attention_mask.sum(dim=1) - 1
        batch_size = last_hidden_states.shape[0]
        return last_hidden_states[torch.arange(batch_size, device=last_hidden_states.device), sequence_lengths]


def get_detailed_instruct(task_description: str, query: str) -> str:
    return f'<instruct>{task_description}\n<query>{query}'


task = 'Given a web search query, retrieve relevant passages that answer the query.'
queries = [
    get_detailed_instruct(task, 'how much protein should a female eat'),
    get_detailed_instruct(task, 'summit define')
]
# No need to add instructions for documents
documents = [
    "As a general guideline, the CDC's average requirement of protein for women ages 19 to 70 is 46 grams per day. But, as you can see from this chart, you'll need to increase that if you're expecting or training for a marathon. Check out the chart below to see how much protein you should be eating each day.",
    "Definition of summit for English Language Learners. : 1  the highest point of a mountain : the top of a mountain. : 2  the highest level. : 3  a meeting or series of meetings between the leaders of two or more governments."
]
input_texts = queries + documents

tokenizer = AutoTokenizer.from_pretrained('BAAI/bge-multilingual-gemma2')
model = AutoModel.from_pretrained('BAAI/bge-multilingual-gemma2')
model.eval()

max_length = 4096
# Tokenize the input texts
batch_dict = tokenizer(input_texts, max_length=max_length, padding=True, truncation=True, return_tensors='pt', pad_to_multiple_of=8)

with torch.no_grad():
    outputs = model(**batch_dict)
    embeddings = last_token_pool(outputs.last_hidden_state, batch_dict['attention_mask'])
    
# normalize embeddings
embeddings = F.normalize(embeddings, p=2, dim=1)
scores = (embeddings[:2] @ embeddings[2:].T) * 100
print(scores.tolist())
# [[55.92064666748047, 1.6549524068832397], [-0.2698777914047241, 49.95653533935547]]
Evaluation

bge-multilingual-gemma2 exhibits state-of-the-art (SOTA) results on benchmarks like MIRACL, MTEB-pl, and MTEB-fr . It also achieves excellent performance on other major evaluations, including MTEB, C-MTEB and AIR-Bench.

nDCG@10: MIRACL-nDCG@10

Recall@100: MIRACL-Recall@100

MTEB-fr/pl MTEB BEIR C-MTEB

Long-Doc (en, Recall@10): AIR-Bench_Long-Doc

QA (en&zh, nDCG@10): AIR-Bench_QA

Model List

bge is short for BAAI general embedding .

Model Language Description query instruction for retrieval [1]
BAAI/bge-multilingual-gemma2 Multilingual - A LLM-based multilingual embedding model, trained on a diverse range of languages and tasks.
BAAI/bge-en-icl English - A LLM-based dense retriever with in-context learning capabilities can fully leverage the model's potential based on a few shot examples(4096 tokens) Provide instructions and few-shot examples freely based on the given task.
BAAI/bge-m3 Multilingual Inference Fine-tune Multi-Functionality(dense retrieval, sparse retrieval, multi-vector(colbert)), Multi-Linguality, and Multi-Granularity(8192 tokens)
BAAI/llm-embedder English Inference Fine-tune a unified embedding model to support diverse retrieval augmentation needs for LLMs See README
BAAI/bge-reranker-large Chinese and English Inference Fine-tune a cross-encoder model which is more accurate but less efficient [2]
BAAI/bge-reranker-base Chinese and English Inference Fine-tune a cross-encoder model which is more accurate but less efficient [2]
BAAI/bge-large-en-v1.5 English Inference Fine-tune version 1.5 with more reasonable similarity distribution Represent this sentence for searching relevant passages:
BAAI/bge-base-en-v1.5 English Inference Fine-tune version 1.5 with more reasonable similarity distribution Represent this sentence for searching relevant passages:
BAAI/bge-small-en-v1.5 English Inference Fine-tune version 1.5 with more reasonable similarity distribution Represent this sentence for searching relevant passages:
BAAI/bge-large-zh-v1.5 Chinese Inference Fine-tune version 1.5 with more reasonable similarity distribution 为这个句子生成表示以用于检索相关文章:
BAAI/bge-base-zh-v1.5 Chinese Inference Fine-tune version 1.5 with more reasonable similarity distribution 为这个句子生成表示以用于检索相关文章:
BAAI/bge-small-zh-v1.5 Chinese Inference Fine-tune version 1.5 with more reasonable similarity distribution 为这个句子生成表示以用于检索相关文章:
BAAI/bge-large-en English Inference Fine-tune :trophy: rank 1st in MTEB leaderboard Represent this sentence for searching relevant passages:
BAAI/bge-base-en English Inference Fine-tune a base-scale model but with similar ability to bge-large-en Represent this sentence for searching relevant passages:
BAAI/bge-small-en English Inference Fine-tune a small-scale model but with competitive performance Represent this sentence for searching relevant passages:
BAAI/bge-large-zh Chinese Inference Fine-tune :trophy: rank 1st in C-MTEB benchmark 为这个句子生成表示以用于检索相关文章:
BAAI/bge-base-zh Chinese Inference Fine-tune a base-scale model but with similar ability to bge-large-zh 为这个句子生成表示以用于检索相关文章:
BAAI/bge-small-zh Chinese Inference Fine-tune a small-scale model but with competitive performance 为这个句子生成表示以用于检索相关文章:
Citation

If you find this repository useful, please consider giving a star :star: and citation

@misc{bge-m3,
      title={BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation}, 
      author={Jianlv Chen and Shitao Xiao and Peitian Zhang and Kun Luo and Defu Lian and Zheng Liu},
      year={2024},
      eprint={2402.03216},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}


@misc{bge_embedding,
      title={C-Pack: Packaged Resources To Advance General Chinese Embedding}, 
      author={Shitao Xiao and Zheng Liu and Peitian Zhang and Niklas Muennighoff},
      year={2023},
      eprint={2309.07597},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}

Runs of BAAI bge-multilingual-gemma2 on huggingface.co

307.7K
Total runs
4.6K
24-hour runs
1.5K
3-day runs
23.3K
7-day runs
-25.5K
30-day runs

More Information About bge-multilingual-gemma2 huggingface.co Model

More bge-multilingual-gemma2 license Visit here:

https://choosealicense.com/licenses/gemma

bge-multilingual-gemma2 huggingface.co

bge-multilingual-gemma2 huggingface.co is an AI model on huggingface.co that provides bge-multilingual-gemma2's model effect (), which can be used instantly with this BAAI bge-multilingual-gemma2 model. huggingface.co supports a free trial of the bge-multilingual-gemma2 model, and also provides paid use of the bge-multilingual-gemma2. Support call bge-multilingual-gemma2 model through api, including Node.js, Python, http.

bge-multilingual-gemma2 huggingface.co Url

https://huggingface.co/BAAI/bge-multilingual-gemma2

BAAI bge-multilingual-gemma2 online free

bge-multilingual-gemma2 huggingface.co is an online trial and call api platform, which integrates bge-multilingual-gemma2's modeling effects, including api services, and provides a free online trial of bge-multilingual-gemma2, you can try bge-multilingual-gemma2 online for free by clicking the link below.

BAAI bge-multilingual-gemma2 online free url in huggingface.co:

https://huggingface.co/BAAI/bge-multilingual-gemma2

bge-multilingual-gemma2 install

bge-multilingual-gemma2 is an open source model from GitHub that offers a free installation service, and any user can find bge-multilingual-gemma2 on GitHub to install. At the same time, huggingface.co provides the effect of bge-multilingual-gemma2 install, users can directly use bge-multilingual-gemma2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

bge-multilingual-gemma2 install url in huggingface.co:

https://huggingface.co/BAAI/bge-multilingual-gemma2

Url of bge-multilingual-gemma2

bge-multilingual-gemma2 huggingface.co Url

Provider of bge-multilingual-gemma2 huggingface.co

BAAI
ORGANIZATIONS

Other API from BAAI

huggingface.co

Total runs: 63.3M
Run Growth: -5.5M
Growth Rate: -8.63%
Updated:February 22 2024
huggingface.co

Total runs: 35.9M
Run Growth: -1.2M
Growth Rate: -3.31%
Updated:July 03 2024
huggingface.co

Total runs: 10.2M
Run Growth: -1.8M
Growth Rate: -18.09%
Updated:February 21 2024
huggingface.co

Total runs: 9.8M
Run Growth: -4.0M
Growth Rate: -40.47%
Updated:February 21 2024
huggingface.co

Total runs: 1.9M
Run Growth: 433.7K
Growth Rate: 22.38%
Updated:April 17 2024
huggingface.co

Total runs: 815.1K
Run Growth: -412.0K
Growth Rate: -50.54%
Updated:April 02 2024
huggingface.co

Total runs: 677.9K
Run Growth: -445.2K
Growth Rate: -65.67%
Updated:October 12 2023
huggingface.co

Total runs: 204.7K
Run Growth: -619.8K
Growth Rate: -302.78%
Updated:December 13 2023
huggingface.co

Total runs: 75.1K
Run Growth: 70.5K
Growth Rate: 93.76%
Updated:October 12 2023
huggingface.co

Total runs: 60.4K
Run Growth: -15.0K
Growth Rate: -24.82%
Updated:July 13 2026
huggingface.co

Total runs: 52.0K
Run Growth: 2.1K
Growth Rate: 4.00%
Updated:October 12 2023
huggingface.co

Total runs: 46.3K
Run Growth: -39.0K
Growth Rate: -84.15%
Updated:October 12 2023
huggingface.co

Total runs: 41.1K
Run Growth: -1.1K
Growth Rate: -2.70%
Updated:January 15 2025
huggingface.co

Total runs: 40.0K
Run Growth: 391
Growth Rate: 0.98%
Updated:November 14 2023
huggingface.co

Total runs: 15.9K
Run Growth: -529
Growth Rate: -3.32%
Updated:April 16 2025
huggingface.co

Total runs: 10.4K
Run Growth: 10.4K
Growth Rate: 100.00%
Updated:September 16 2026
huggingface.co

Total runs: 9.1K
Run Growth: -10.6K
Growth Rate: -117.46%
Updated:October 12 2023
huggingface.co

Total runs: 5.8K
Run Growth: 299
Growth Rate: 5.20%
Updated:May 20 2025
huggingface.co

Total runs: 4.7K
Run Growth: 3.2K
Growth Rate: 68.88%
Updated:April 19 2024
huggingface.co

Total runs: 2.4K
Run Growth: 765
Growth Rate: 31.77%
Updated:April 10 2026
huggingface.co

Total runs: 2.4K
Run Growth: 330
Growth Rate: 13.90%
Updated:September 11 2026
huggingface.co

Total runs: 1.8K
Run Growth: 953
Growth Rate: 52.56%
Updated:November 28 2024
huggingface.co

Total runs: 1.8K
Run Growth: -13.6K
Growth Rate: -739.70%
Updated:January 15 2025
huggingface.co

Total runs: 1.7K
Run Growth: 326
Growth Rate: 18.87%
Updated:April 10 2026
huggingface.co

Total runs: 1.6K
Run Growth: 395
Growth Rate: 24.28%
Updated:December 31 2022
huggingface.co

Total runs: 1.3K
Run Growth: 849
Growth Rate: 64.56%
Updated:September 11 2026
huggingface.co

Total runs: 727
Run Growth: -85
Growth Rate: -11.69%
Updated:May 23 2025
huggingface.co

Total runs: 706
Run Growth: -850
Growth Rate: -121.78%
Updated:February 07 2024
huggingface.co

Total runs: 669
Run Growth: -90
Growth Rate: -13.45%
Updated:October 24 2024
huggingface.co

Total runs: 665
Run Growth: -440
Growth Rate: -66.17%
Updated:October 23 2024
huggingface.co

Total runs: 514
Run Growth: -358
Growth Rate: -69.65%
Updated:April 18 2023
huggingface.co

Total runs: 476
Run Growth: -249
Growth Rate: -52.31%
Updated:July 13 2026
huggingface.co

Total runs: 437
Run Growth: 23
Growth Rate: 5.26%
Updated:April 10 2026
huggingface.co

Total runs: 429
Run Growth: -99
Growth Rate: -23.08%
Updated:August 11 2026
huggingface.co

Total runs: 400
Run Growth: 246
Growth Rate: 61.50%
Updated:April 02 2024
huggingface.co

Total runs: 388
Run Growth: 46
Growth Rate: 11.86%
Updated:July 13 2026
huggingface.co

Total runs: 384
Run Growth: -15
Growth Rate: -3.91%
Updated:December 25 2025
huggingface.co

Total runs: 375
Run Growth: -373
Growth Rate: -99.47%
Updated:August 11 2026
huggingface.co

Total runs: 370
Run Growth: -200
Growth Rate: -54.05%
Updated:July 13 2026
huggingface.co

Total runs: 304
Run Growth: 56
Growth Rate: 18.42%
Updated:July 13 2026
huggingface.co

Total runs: 293
Run Growth: -124
Growth Rate: -42.32%
Updated:July 13 2026
huggingface.co

Total runs: 291
Run Growth: -345
Growth Rate: -117.35%
Updated:December 25 2025
huggingface.co

Total runs: 273
Run Growth: -183
Growth Rate: -67.03%
Updated:October 27 2023
huggingface.co

Total runs: 272
Run Growth: -1
Growth Rate: -0.37%
Updated:July 13 2026
huggingface.co

Total runs: 270
Run Growth: -196
Growth Rate: -78.71%
Updated:March 07 2024
huggingface.co

Total runs: 265
Run Growth: 52
Growth Rate: 19.62%
Updated:August 18 2026
huggingface.co

Total runs: 249
Run Growth: 249
Growth Rate: 100.00%
Updated:September 05 2026
huggingface.co

Total runs: 245
Run Growth: 16
Growth Rate: 6.53%
Updated:July 13 2026
huggingface.co

Total runs: 229
Run Growth: -17
Growth Rate: -7.42%
Updated:June 06 2025
huggingface.co

Total runs: 223
Run Growth: 19
Growth Rate: 8.52%
Updated:July 17 2026
huggingface.co

Total runs: 220
Run Growth: -2.0K
Growth Rate: -888.64%
Updated:April 10 2026
huggingface.co

Total runs: 215
Run Growth: 20
Growth Rate: 9.30%
Updated:July 13 2026
huggingface.co

Total runs: 212
Run Growth: 122
Growth Rate: 57.55%
Updated:November 28 2023
huggingface.co

Total runs: 210
Run Growth: -168
Growth Rate: -80.00%
Updated:April 10 2026
huggingface.co

Total runs: 199
Run Growth: 90
Growth Rate: 45.23%
Updated:July 13 2026
huggingface.co

Total runs: 199
Run Growth: 86
Growth Rate: 43.22%
Updated:July 13 2026
huggingface.co

Total runs: 192
Run Growth: -18
Growth Rate: -9.38%
Updated:October 23 2024
huggingface.co

Total runs: 191
Run Growth: 97
Growth Rate: 50.79%
Updated:March 13 2026
huggingface.co

Total runs: 180
Run Growth: 29
Growth Rate: 16.11%
Updated:June 24 2024
huggingface.co

Total runs: 178
Run Growth: -24
Growth Rate: -13.48%
Updated:May 13 2024
huggingface.co

Total runs: 177
Run Growth: 1
Growth Rate: 0.56%
Updated:May 13 2024
huggingface.co

Total runs: 177
Run Growth: 28
Growth Rate: 15.82%
Updated:June 24 2024
huggingface.co

Total runs: 176
Run Growth: 77
Growth Rate: 43.75%
Updated:December 21 2023
huggingface.co

Total runs: 173
Run Growth: -2
Growth Rate: -1.16%
Updated:May 13 2024