Wings is a brand-new universal Multimodal Large Language Model (MLLM). Its flexible multimodal structure enhances the MLLM as if
giving it wings that enhance the performance of multimodal capabilities
while minimizing text-only forgetting.
Any
architecture of MLLM can adapt the Wings component.
Multimodal large language models (MLLMs), initiated with a trained LLM, first align images with text and then fine-tune on multimodal mixed inputs. However, the MLLM
catastrophically forgets the text-only instructions
, which do not include images and can be addressed within the initial LLM.
In this work, we present Wings, a novel MLLM that excels in both text-only dialogues and multimodal comprehension. Analyzing MLLM attention in multimodal instructions reveals that
text-only forgetting is related to the attention shifts from pre-image to post-image text.
From that, we construct extra modules that act as the boosted learner to compensate for the attention shift. The complementary visual and textual learners,
like "wings" on either side, are connected in parallel within each layer's attention block.
Initially, image and text inputs are aligned with visual learners operating alongside the main attention, balancing focus on visual elements. Textual learners are later collaboratively integrated with attention-based routing to blend the outputs of the visual and textual learners. We design the
Low-Rank Residual Attention (LoRRA)
to guarantee high efficiency for learners.
Our experimental results demonstrate that Wings outperforms equally-scaled MLLMs in both text-only and visual question-answering tasks. On a newly constructed Interleaved Image-Text (IIT) benchmark, Wings exhibits superior performance from text-only-rich to multimodal-rich question-answering tasks.
bash run/pretrain_base.sh
# Set path for pretrained MLLM
bash run/finetune_base.sh
Citation
If you find Wings useful, please cite the paper:
@article{zhang_wings,
author = {Yi{-}Kai Zhang and
Shiyin Lu and
Yang Li and
Yanqing Ma and
Qing{-}Guo Chen and
Zhao Xu and
Weihua Luo and
Kaifu Zhang and
De{-}Chuan Zhan and
Han{-}Jia Ye},
title = {Wings: Learning Multimodal LLMs without Text-only Forgetting},
journal = {CoRR},
volume = {abs/2406.03496},
year = {2024}
}
We used compliance-checking algorithms during the training process, to ensure the compliance of the trained model to the best of our ability. Due to the complexity of the data and the diversity of language model usage scenarios, we cannot guarantee that the model is completely free of copyright issues or improper content. If you believe anything infringes on your rights or generates improper content, please contact us, and we will promptly address the matter.
Runs of AIDC-AI Wings-Qwen1_5-8B on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About Wings-Qwen1_5-8B huggingface.co Model
Wings-Qwen1_5-8B huggingface.co is an AI model on huggingface.co that provides Wings-Qwen1_5-8B's model effect (), which can be used instantly with this AIDC-AI Wings-Qwen1_5-8B model. huggingface.co supports a free trial of the Wings-Qwen1_5-8B model, and also provides paid use of the Wings-Qwen1_5-8B. Support call Wings-Qwen1_5-8B model through api, including Node.js, Python, http.
Wings-Qwen1_5-8B huggingface.co is an online trial and call api platform, which integrates Wings-Qwen1_5-8B's modeling effects, including api services, and provides a free online trial of Wings-Qwen1_5-8B, you can try Wings-Qwen1_5-8B online for free by clicking the link below.
AIDC-AI Wings-Qwen1_5-8B online free url in huggingface.co:
Wings-Qwen1_5-8B is an open source model from GitHub that offers a free installation service, and any user can find Wings-Qwen1_5-8B on GitHub to install. At the same time, huggingface.co provides the effect of Wings-Qwen1_5-8B install, users can directly use Wings-Qwen1_5-8B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.