Based on pre-trained clip-vit-base (patch 16, resolution 224x224), we apply some multimodal information with special pre-training tasks. "D" implies a special training method. For special multimodal representations, we design several special training objectives in our paper. The pre-training datasets are MSCOCO and VG. Our code and details of pre-training tasks will be made publicly available upon paper acceptance.
下游任务 Performance
CIFAR10
ImageNet1k
clip-vit-base-patch16-224 (official)
96.2
80.2
Taiyi-vit-87M-D (local)
98.7
82.4
The local test settings are:
learning rate=2e-5,
batch size=128,
num train epochs=5,
weight decay=0.01
使用 Usage
from transformers import ViTFeatureExtractor, ViTForImageClassification
from PIL import Image
import requests
url = 'http://images.cocodataset.org/val2017/000000039769.jpg'
image = Image.open(requests.get(url, stream=True).raw)
feature_extractor = ViTFeatureExtractor.from_pretrained('IDEA-CCNL/Taiyi-vit-87M-D')
model = ViTForImageClassification.from_pretrained('IDEA-CCNL/Taiyi-vit-87M-D')
inputs = feature_extractor(images=image, return_tensors="pt")
outputs = model(**inputs)
logits = outputs.logits
# model predicts one of the 1000 ImageNet classes
predicted_class_idx = logits.argmax(-1).item()
print("Predicted class:", model.config.id2label[predicted_class_idx])
# Predicted class: Egyptian cat
If you are using the resource for your work, please cite the our
paper
:
@article{fengshenbang,
author = {Jiaxing Zhang and Ruyi Gan and Junjie Wang and Yuxiang Zhang and Lin Zhang and Ping Yang and Xinyu Gao and Ziwei Wu and Xiaoqun Dong and Junqing He and Jianheng Zhuo and Qi Yang and Yongfeng Huang and Xiayu Li and Yanghan Wu and Junyu Lu and Xinyu Zhu and Weifeng Chen and Ting Han and Kunhao Pan and Rui Wang and Hao Wang and Xiaojun Wu and Zhongshen Zeng and Chongpei Chen},
title = {Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence},
journal = {CoRR},
volume = {abs/2209.02970},
year = {2022}
}
Taiyi-vit-87M-D huggingface.co is an AI model on huggingface.co that provides Taiyi-vit-87M-D's model effect (), which can be used instantly with this IDEA-CCNL Taiyi-vit-87M-D model. huggingface.co supports a free trial of the Taiyi-vit-87M-D model, and also provides paid use of the Taiyi-vit-87M-D. Support call Taiyi-vit-87M-D model through api, including Node.js, Python, http.
Taiyi-vit-87M-D huggingface.co is an online trial and call api platform, which integrates Taiyi-vit-87M-D's modeling effects, including api services, and provides a free online trial of Taiyi-vit-87M-D, you can try Taiyi-vit-87M-D online for free by clicking the link below.
IDEA-CCNL Taiyi-vit-87M-D online free url in huggingface.co:
Taiyi-vit-87M-D is an open source model from GitHub that offers a free installation service, and any user can find Taiyi-vit-87M-D on GitHub to install. At the same time, huggingface.co provides the effect of Taiyi-vit-87M-D install, users can directly use Taiyi-vit-87M-D installed effect in huggingface.co for debugging and trial. It also supports api for free installation.