Salesforce / blip-image-captioning-base

huggingface.co
Total runs: 1.9M
24-hour runs: 0
7-day runs: -139.0K
30-day runs: 152.4K
Model's Last Updated: February 03 2025
image-to-text

Introduction of blip-image-captioning-base

Model Details of blip-image-captioning-base

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Model card for image captioning pretrained on COCO dataset - base architecture (with ViT base backbone).

BLIP.gif
Pull figure from BLIP official repo
TL;DR

Authors from the paper write in the abstract:

Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to videolanguage tasks in a zero-shot manner. Code, models, and datasets are released.

Usage

You can use this model for conditional and un-conditional image captioning

Using the Pytorch model
Running the model on CPU
Click to expand
import requests
from PIL import Image
from transformers import BlipProcessor, BlipForConditionalGeneration

processor = BlipProcessor.from_pretrained("Salesforce/blip-image-captioning-base")
model = BlipForConditionalGeneration.from_pretrained("Salesforce/blip-image-captioning-base")

img_url = 'https://storage.googleapis.com/sfr-vision-language-research/BLIP/demo.jpg' 
raw_image = Image.open(requests.get(img_url, stream=True).raw).convert('RGB')

# conditional image captioning
text = "a photography of"
inputs = processor(raw_image, text, return_tensors="pt")

out = model.generate(**inputs)
print(processor.decode(out[0], skip_special_tokens=True))
# >>> a photography of a woman and her dog

# unconditional image captioning
inputs = processor(raw_image, return_tensors="pt")

out = model.generate(**inputs)
print(processor.decode(out[0], skip_special_tokens=True))
>>> a woman sitting on the beach with her dog
Running the model on GPU
In full precision
Click to expand
import requests
from PIL import Image
from transformers import BlipProcessor, BlipForConditionalGeneration

processor = BlipProcessor.from_pretrained("Salesforce/blip-image-captioning-base")
model = BlipForConditionalGeneration.from_pretrained("Salesforce/blip-image-captioning-base").to("cuda")

img_url = 'https://storage.googleapis.com/sfr-vision-language-research/BLIP/demo.jpg' 
raw_image = Image.open(requests.get(img_url, stream=True).raw).convert('RGB')

# conditional image captioning
text = "a photography of"
inputs = processor(raw_image, text, return_tensors="pt").to("cuda")

out = model.generate(**inputs)
print(processor.decode(out[0], skip_special_tokens=True))
# >>> a photography of a woman and her dog

# unconditional image captioning
inputs = processor(raw_image, return_tensors="pt").to("cuda")

out = model.generate(**inputs)
print(processor.decode(out[0], skip_special_tokens=True))
>>> a woman sitting on the beach with her dog
In half precision ( float16 )
Click to expand
import torch
import requests
from PIL import Image
from transformers import BlipProcessor, BlipForConditionalGeneration

processor = BlipProcessor.from_pretrained("Salesforce/blip-image-captioning-base")
model = BlipForConditionalGeneration.from_pretrained("Salesforce/blip-image-captioning-base", torch_dtype=torch.float16).to("cuda")

img_url = 'https://storage.googleapis.com/sfr-vision-language-research/BLIP/demo.jpg' 
raw_image = Image.open(requests.get(img_url, stream=True).raw).convert('RGB')

# conditional image captioning
text = "a photography of"
inputs = processor(raw_image, text, return_tensors="pt").to("cuda", torch.float16)

out = model.generate(**inputs)
print(processor.decode(out[0], skip_special_tokens=True))
# >>> a photography of a woman and her dog

# unconditional image captioning
inputs = processor(raw_image, return_tensors="pt").to("cuda", torch.float16)

out = model.generate(**inputs)
print(processor.decode(out[0], skip_special_tokens=True))
>>> a woman sitting on the beach with her dog
BibTex and citation info
@misc{https://doi.org/10.48550/arxiv.2201.12086,
  doi = {10.48550/ARXIV.2201.12086},
  
  url = {https://arxiv.org/abs/2201.12086},
  
  author = {Li, Junnan and Li, Dongxu and Xiong, Caiming and Hoi, Steven},
  
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  
  title = {BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation},
  
  publisher = {arXiv},
  
  year = {2022},
  
  copyright = {Creative Commons Attribution 4.0 International}
}

Runs of Salesforce blip-image-captioning-base on huggingface.co

1.9M
Total runs
0
24-hour runs
-68.5K
3-day runs
-139.0K
7-day runs
152.4K
30-day runs

More Information About blip-image-captioning-base huggingface.co Model

More blip-image-captioning-base license Visit here:

https://choosealicense.com/licenses/bsd-3-clause

blip-image-captioning-base huggingface.co

blip-image-captioning-base huggingface.co is an AI model on huggingface.co that provides blip-image-captioning-base's model effect (), which can be used instantly with this Salesforce blip-image-captioning-base model. huggingface.co supports a free trial of the blip-image-captioning-base model, and also provides paid use of the blip-image-captioning-base. Support call blip-image-captioning-base model through api, including Node.js, Python, http.

blip-image-captioning-base huggingface.co Url

https://huggingface.co/Salesforce/blip-image-captioning-base

Salesforce blip-image-captioning-base online free

blip-image-captioning-base huggingface.co is an online trial and call api platform, which integrates blip-image-captioning-base's modeling effects, including api services, and provides a free online trial of blip-image-captioning-base, you can try blip-image-captioning-base online for free by clicking the link below.

Salesforce blip-image-captioning-base online free url in huggingface.co:

https://huggingface.co/Salesforce/blip-image-captioning-base

blip-image-captioning-base install

blip-image-captioning-base is an open source model from GitHub that offers a free installation service, and any user can find blip-image-captioning-base on GitHub to install. At the same time, huggingface.co provides the effect of blip-image-captioning-base install, users can directly use blip-image-captioning-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

blip-image-captioning-base install url in huggingface.co:

https://huggingface.co/Salesforce/blip-image-captioning-base

Url of blip-image-captioning-base

blip-image-captioning-base huggingface.co Url

Provider of blip-image-captioning-base huggingface.co

Salesforce
ORGANIZATIONS

Other API from Salesforce

huggingface.co

Total runs: 98.3K
Run Growth: -2.9K
Growth Rate: -2.96%
Updated:February 03 2025
huggingface.co

Total runs: 80.1K
Run Growth: 80.0K
Growth Rate: 99.89%
Updated:April 12 2025
huggingface.co

Total runs: 250
Run Growth: -278
Growth Rate: -111.20%
Updated:January 15 2025
huggingface.co

Total runs: 242
Run Growth: -5.1K
Growth Rate: -2112.40%
Updated:October 04 2025
huggingface.co

Total runs: 185
Run Growth: -60
Growth Rate: -32.43%
Updated:January 21 2025
huggingface.co

Total runs: 150
Run Growth: -3
Growth Rate: -2.00%
Updated:November 05 2025