This work proposes
Multimodal Adversarial Training (MAT)
for Vision-Language Models (VLMs). MAT is a unified adversarial training pipeline for image-text retrieval models. The extended version,
MAT+
, additionally leverages one-to-many relationships in image-text pairs to improve robustness.
Highlights
Unified MAT pipeline for image-text retrieval models (CLIP, ALBEF, BLIP).
MAT+ leverages one-to-many relationships in image-text pairs.
Reproducible results on Flickr30k and COCO benchmarks.
Data augmentations used to reproduce MAT+ results:
File
Description
dataset_json.zip
Text augmentation data โ augmented captions and annotations in JSON format
flickr_SD_I2I_0.5.zip
Image augmentation data โ Flickr30k images augmented via Stable Diffusion image-to-image (strength 0.5)
๐ Usage
Clone or download this repository:
# Using the Hugging Face CLI
hf download cyberagent/multimodal-adversarial-training --local-dir ./resources
# Or using git with LFS
git lfs install
git clone https://huggingface.co/cyberagent/multimodal-adversarial-training
Update the checkpoint and data paths in
configs/
to point to the downloaded resources.
๐ Citation
If you find these resources useful, please cite:
@inproceedings{waseda2026multimodal,
title={Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships},
author={Waseda, Futa and Tejero-de-Pablos, Antonio and Echizen, Isao},
booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
year={2026}
}
multimodal-adversarial-training huggingface.co is an AI model on huggingface.co that provides multimodal-adversarial-training's model effect (), which can be used instantly with this cyberagent multimodal-adversarial-training model. huggingface.co supports a free trial of the multimodal-adversarial-training model, and also provides paid use of the multimodal-adversarial-training. Support call multimodal-adversarial-training model through api, including Node.js, Python, http.
multimodal-adversarial-training huggingface.co is an online trial and call api platform, which integrates multimodal-adversarial-training's modeling effects, including api services, and provides a free online trial of multimodal-adversarial-training, you can try multimodal-adversarial-training online for free by clicking the link below.
cyberagent multimodal-adversarial-training online free url in huggingface.co:
multimodal-adversarial-training is an open source model from GitHub that offers a free installation service, and any user can find multimodal-adversarial-training on GitHub to install. At the same time, huggingface.co provides the effect of multimodal-adversarial-training install, users can directly use multimodal-adversarial-training installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
multimodal-adversarial-training install url in huggingface.co: