import torch
import torch.nn.functional as F
from urllib.request import urlopen
from PIL import Image
from open_clip import create_model_from_pretrained, get_tokenizer # works on open-clip-torch >= 2.31.0, timm >= 1.0.15
model, preprocess = create_model_from_pretrained('hf-hub:timm/ViT-B-16-SigLIP2-256')
tokenizer = get_tokenizer('hf-hub:timm/ViT-B-16-SigLIP2-256')
image = Image.open(urlopen(
'https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/beignets-task-guide.png'
))
image = preprocess(image).unsqueeze(0)
labels_list = ["a dog", "a cat", "a donut", "a beignet"]
text = tokenizer(labels_list, context_length=model.context_length)
with torch.no_grad(), torch.cuda.amp.autocast():
image_features = model.encode_image(image, normalize=True)
text_features = model.encode_text(text, normalize=True)
text_probs = torch.sigmoid(image_features @ text_features.T * model.logit_scale.exp() + model.logit_bias)
zipped_list = list(zip(labels_list, [100 * round(p.item(), 3) for p in text_probs[0]]))
print("Label probabilities: ", zipped_list)
Citation
@article{tschannen2025siglip,
title={SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features},
author={Tschannen, Michael and Gritsenko, Alexey and Wang, Xiao and Naeem, Muhammad Ferjad and Alabdulmohsin, Ibrahim and Parthasarathy, Nikhil and Evans, Talfan and Beyer, Lucas and Xia, Ye and Mustafa, Basil and H'enaff, Olivier and Harmsen, Jeremiah and Steiner, Andreas and Zhai, Xiaohua},
year={2025},
journal={arXiv preprint arXiv:2502.14786}
}
@article{zhai2023sigmoid,
title={Sigmoid loss for language image pre-training},
author={Zhai, Xiaohua and Mustafa, Basil and Kolesnikov, Alexander and Beyer, Lucas},
journal={arXiv preprint arXiv:2303.15343},
year={2023}
}
@misc{big_vision,
author = {Beyer, Lucas and Zhai, Xiaohua and Kolesnikov, Alexander},
title = {Big Vision},
year = {2022},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/google-research/big_vision}}
}
Runs of timm ViT-B-16-SigLIP2-256 on huggingface.co
131.4K
Total runs
0
24-hour runs
3.7K
3-day runs
3.7K
7-day runs
-54.1K
30-day runs
More Information About ViT-B-16-SigLIP2-256 huggingface.co Model
ViT-B-16-SigLIP2-256 huggingface.co is an AI model on huggingface.co that provides ViT-B-16-SigLIP2-256's model effect (), which can be used instantly with this timm ViT-B-16-SigLIP2-256 model. huggingface.co supports a free trial of the ViT-B-16-SigLIP2-256 model, and also provides paid use of the ViT-B-16-SigLIP2-256. Support call ViT-B-16-SigLIP2-256 model through api, including Node.js, Python, http.
ViT-B-16-SigLIP2-256 huggingface.co is an online trial and call api platform, which integrates ViT-B-16-SigLIP2-256's modeling effects, including api services, and provides a free online trial of ViT-B-16-SigLIP2-256, you can try ViT-B-16-SigLIP2-256 online for free by clicking the link below.
timm ViT-B-16-SigLIP2-256 online free url in huggingface.co:
ViT-B-16-SigLIP2-256 is an open source model from GitHub that offers a free installation service, and any user can find ViT-B-16-SigLIP2-256 on GitHub to install. At the same time, huggingface.co provides the effect of ViT-B-16-SigLIP2-256 install, users can directly use ViT-B-16-SigLIP2-256 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
ViT-B-16-SigLIP2-256 install url in huggingface.co: