DeepGlint-AI / UniME-LLaVA-OneVision-7B

huggingface.co
Total runs: 1.6K
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: May 07 2025
image-text-to-text

Introduction of UniME-LLaVA-OneVision-7B

Model Details of UniME-LLaVA-OneVision-7B

Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs

Tiancheng Gu* , Kaicheng Yang* , Ziyong Feng, Xingjun Wang, Yanzhao Zhang, Dingkun Long, Yingda Chen, Weidong Cai , Jiankang Deng

🏡 Project Page | 📄 Paper | 💻 Github

UniME achieves the top ranking on the MMEB leaderboard training using a 336×336 image resolution.(The screenshot is captured at 08:00 UTC+8 on May 6, 2025.)

💡 Highlights

To enhance the MLLM's embedding capability, we propose textual discriminative knowledge distillation. The training process involves decoupling the MLLM's LLM component and processing text with the prompt "Summarize the above sentences in one word.", followed by aligning the student (MLLM) and teacher (NV-Embed V2) embeddings via KL divergence on batch-wise similarity distributions. Notably, only the LLM component is fine-tuned during this process, while all other parameters remain frozen .

After that, we propose hard negative enhanced instruction tuning enhances multimodal systems by improving visual sensitivity, strengthening cross-modal alignment, and boosting instruction-following capabilities. At its core are two key innovations: a false negative filtering mechanism using a similarity threshold to eliminate misleading samples, and an automatic hard negative sampling strategy that selects top-k similar but non-matching examples to increase training difficulty.

🧭 Quick Start
git clone https://github.com/deepglint/UniME.git
cd UniME
conda create -n uniME python=3.10 -y
conda activate uniME
pip install -r requirements.txt
pip install transformers==4.49.0
import torch
from PIL import Image
from torch.nn import functional as F
from transformers import AutoProcessor, LlavaOnevisionForConditionalGeneration

def appply_chat_template(image=None, text=None):
    if image != None:
        conversation_image = [{
                "role": "user",
                "content": [
                    {"type": "image", "image": image},
                    {"type": "text", "text": "Summary above image in one word:\n"},
                    ],
            }]
    elif text!= None:
        conversation_image = [{
                "role": "user",
                "content": [
                    {"type": "text", "text": f"{text}\nSummary above sentence in one word:\n"},
                    ],
            }]
    return conversation_image

base_model_path="DeepGlint-AI/UniME-LLaVA-OneVision-7B"

text = "A man is crossing the street with a red car parked nearby."
image_path = "figures/demo.png"
input_image = [Image.open(image_path)]

transform = AutoProcessor.from_pretrained(base_model_path, trust_remote_code=True)
model = LlavaOnevisionForConditionalGeneration.from_pretrained(base_model_path,device_map="cuda", trust_remote_code=True, torch_dtype=torch.float16)
transform.tokenizer.padding_side = "left"
transform.tokenizer.padding = True

inputs_text = transform.apply_chat_template([appply_chat_template(text = text)],
                                        add_generation_prompt=True,
                                        tokenize=True,
                                        return_dict=True,
                                        return_tensors="pt",
                                        padding=True).to("cuda")
inputs_image = transform.apply_chat_template([appply_chat_template(image = input_image)],
                                        add_generation_prompt=True,
                                        tokenize=True,
                                        return_dict=True,
                                        return_tensors="pt",
                                        padding=True).to("cuda")

with torch.no_grad():
  emb_text = model(**inputs_text, output_hidden_states=True, return_dict=True).hidden_states[-1][:, -1, :]
  emb_image = model(**inputs_image, output_hidden_states=True, return_dict=True).hidden_states[-1][:, -1, :]
  emb_text = F.normalize(emb_text, dim=-1)
  emb_image = F.normalize(emb_image, dim=-1)
  Score = emb_image @ emb_text.T
print("Score: ", Score.item())
🔢 Results
Diverse Retrieval

MMEB

📖 Citation

If you find this repository useful, please use the following BibTeX entry for citation.

@misc{gu2025breakingmodalitybarrieruniversal,
      title={Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs}, 
      author={Tiancheng Gu and Kaicheng Yang and Ziyong Feng and Xingjun Wang and Yanzhao Zhang and Dingkun Long and Yingda Chen and Weidong Cai and Jiankang Deng},
      year={2025},
      eprint={2504.17432},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2504.17432}, 
}

Runs of DeepGlint-AI UniME-LLaVA-OneVision-7B on huggingface.co

1.6K
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About UniME-LLaVA-OneVision-7B huggingface.co Model

More UniME-LLaVA-OneVision-7B license Visit here:

https://choosealicense.com/licenses/mit

UniME-LLaVA-OneVision-7B huggingface.co

UniME-LLaVA-OneVision-7B huggingface.co is an AI model on huggingface.co that provides UniME-LLaVA-OneVision-7B's model effect (), which can be used instantly with this DeepGlint-AI UniME-LLaVA-OneVision-7B model. huggingface.co supports a free trial of the UniME-LLaVA-OneVision-7B model, and also provides paid use of the UniME-LLaVA-OneVision-7B. Support call UniME-LLaVA-OneVision-7B model through api, including Node.js, Python, http.

UniME-LLaVA-OneVision-7B huggingface.co Url

https://huggingface.co/DeepGlint-AI/UniME-LLaVA-OneVision-7B

DeepGlint-AI UniME-LLaVA-OneVision-7B online free

UniME-LLaVA-OneVision-7B huggingface.co is an online trial and call api platform, which integrates UniME-LLaVA-OneVision-7B's modeling effects, including api services, and provides a free online trial of UniME-LLaVA-OneVision-7B, you can try UniME-LLaVA-OneVision-7B online for free by clicking the link below.

DeepGlint-AI UniME-LLaVA-OneVision-7B online free url in huggingface.co:

https://huggingface.co/DeepGlint-AI/UniME-LLaVA-OneVision-7B

UniME-LLaVA-OneVision-7B install

UniME-LLaVA-OneVision-7B is an open source model from GitHub that offers a free installation service, and any user can find UniME-LLaVA-OneVision-7B on GitHub to install. At the same time, huggingface.co provides the effect of UniME-LLaVA-OneVision-7B install, users can directly use UniME-LLaVA-OneVision-7B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

UniME-LLaVA-OneVision-7B install url in huggingface.co:

https://huggingface.co/DeepGlint-AI/UniME-LLaVA-OneVision-7B

Url of UniME-LLaVA-OneVision-7B

UniME-LLaVA-OneVision-7B huggingface.co Url

Provider of UniME-LLaVA-OneVision-7B huggingface.co

DeepGlint-AI
ORGANIZATIONS

Other API from DeepGlint-AI