BLIP-2 (Bootstrapping Language-Image Pre-training) is a generic and efficient pre-training strategy that bridges the modality gap between frozen image encoders and frozen large language models (LLMs). It was introduced by Salesforce in the paper "BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models".
The architecture consists of three main components: a frozen image encoder (EVA-CLIP ViT-g/14), a lightweight querying transformer (Q-Former) that acts as an information bottleneck to extract the most relevant visual features, and a frozen LLM (like OPT or Flan-T5) that handles the text generation. Because the heavy vision and language models are kept frozen during pre-training, BLIP-2 achieves state-of-the-art performance on various vision-language tasks with significantly fewer trainable parameters than existing methods.
pip install -U -q keras-hub
pip install -U -q keras
Jax, TensorFlow, and Torch come preinstalled in Kaggle Notebooks. For instructions on installing them in another environment see the
Keras Getting Started
page.
Preset Table
Preset
Architecture
Vision Encoder
Language Model
Description
blip2_opt_2.7b
BLIP-2
EVA-CLIP ViT-g/14
OPT-2.7B
BLIP-2 model using OPT-2.7B as the frozen language model.
blip2_opt_6.7b
BLIP-2
EVA-CLIP ViT-g/14
OPT-6.7B
BLIP-2 model using OPT-6.7B as the frozen language model.
blip2_flan_t5_xl
BLIP-2
EVA-CLIP ViT-g/14
Flan-T5-XL
BLIP-2 model using Flan-T5-XL (~3B) as the frozen language model.
blip2_flan_t5_xxl
BLIP-2
EVA-CLIP ViT-g/14
Flan-T5-XXL
BLIP-2 model using Flan-T5-XXL (~11B) as the frozen language model.
Example Usage
BLIP-2 can be used for various vision-language tasks such as image captioning and visual question answering (VQA). Depending on the preset you choose, you will use either
BLIP2CausalLM
(for OPT models) or
BLIP2Seq2SeqLM
(for Flan-T5 models).
The
BLIP2Seq2SeqLM
class supports BLIP-2 variants that use Flan-T5 as their base language model. This architecture requires passing your text prompt to the encoder using the
"encoder_text"
key.
import keras
import keras_hub
import numpy as np
from PIL import Image
import requests
model = keras_hub.models.BLIP2Seq2SeqLM.from_preset("blip2_opt_6.7b")
image_url = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"
image = Image.open(requests.get(image_url, stream=True).raw).convert("RGB")
image_array = np.array(image)
vqa_input = {
"images": image_array,
"encoder_text": ["Question: what is in the picture? Answer:"]
}
print(model.generate(vqa_input))
caption_input = {
"images": image_array,
"encoder_text": ["A picture of"]
}
print(model.generate(caption_input))
Example Usage with Hugging Face URI
BLIP-2 can be used for various vision-language tasks such as image captioning and visual question answering (VQA). Depending on the preset you choose, you will use either
BLIP2CausalLM
(for OPT models) or
BLIP2Seq2SeqLM
(for Flan-T5 models).
The
BLIP2Seq2SeqLM
class supports BLIP-2 variants that use Flan-T5 as their base language model. This architecture requires passing your text prompt to the encoder using the
"encoder_text"
key.
import keras
import keras_hub
import numpy as np
from PIL import Image
import requests
model = keras_hub.models.BLIP2Seq2SeqLM.from_preset("hf://keras/blip2_opt_6.7b")
image_url = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"
image = Image.open(requests.get(image_url, stream=True).raw).convert("RGB")
image_array = np.array(image)
vqa_input = {
"images": image_array,
"encoder_text": ["Question: what is in the picture? Answer:"]
}
print(model.generate(vqa_input))
caption_input = {
"images": image_array,
"encoder_text": ["A picture of"]
}
print(model.generate(caption_input))
Runs of keras blip2_opt_6.7b on huggingface.co
16
Total runs
-1
24-hour runs
-3
3-day runs
-4
7-day runs
-1
30-day runs
More Information About blip2_opt_6.7b huggingface.co Model
blip2_opt_6.7b huggingface.co
blip2_opt_6.7b huggingface.co is an AI model on huggingface.co that provides blip2_opt_6.7b's model effect (), which can be used instantly with this keras blip2_opt_6.7b model. huggingface.co supports a free trial of the blip2_opt_6.7b model, and also provides paid use of the blip2_opt_6.7b. Support call blip2_opt_6.7b model through api, including Node.js, Python, http.
blip2_opt_6.7b huggingface.co is an online trial and call api platform, which integrates blip2_opt_6.7b's modeling effects, including api services, and provides a free online trial of blip2_opt_6.7b, you can try blip2_opt_6.7b online for free by clicking the link below.
keras blip2_opt_6.7b online free url in huggingface.co:
blip2_opt_6.7b is an open source model from GitHub that offers a free installation service, and any user can find blip2_opt_6.7b on GitHub to install. At the same time, huggingface.co provides the effect of blip2_opt_6.7b install, users can directly use blip2_opt_6.7b installed effect in huggingface.co for debugging and trial. It also supports api for free installation.