cjwbw / clip-vit-large-patch14

openai/clip-vit-large-patch14 with Transformers

replicate.com
Total runs: 6.9M
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Github
Model's Last Updated: September 22 2022

Introduction of clip-vit-large-patch14

Model Details of clip-vit-large-patch14

Readme

openai/clip-vit-large-patch14 with Transformers

Model Card: CLIP

Disclaimer: The model card is taken and modified from the official CLIP repository, it can be found here .

Model Details

The CLIP model was developed by researchers at OpenAI to learn about what contributes to robustness in computer vision tasks. The model was also developed to test the ability of models to generalize to arbitrary image classification tasks in a zero-shot manner. It was not developed for general model deployment - to deploy models like CLIP, researchers will first need to carefully study their capabilities in relation to the specific context they’re being deployed within.

Model Date

January 2021

Model Type

The base model uses a ViT-L/14 Transformer architecture as an image encoder and uses a masked self-attention Transformer as a text encoder. These encoders are trained to maximize the similarity of (image, text) pairs via a contrastive loss.

The original implementation had two variants: one using a ResNet image encoder and the other using a Vision Transformer. This repository has the variant with the Vision Transformer.

Documents
Model Use
Intended Use

The model is intended as a research output for research communities. We hope that this model will enable researchers to better understand and explore zero-shot, arbitrary image classification. We also hope it can be used for interdisciplinary studies of the potential impact of such models - the CLIP paper includes a discussion of potential downstream impacts to provide an example for this sort of analysis.

Primary intended uses

The primary intended users of these models are AI researchers.

We primarily imagine the model will be used by researchers to better understand robustness, generalization, and other capabilities, biases, and constraints of computer vision models.

Out-of-Scope Use Cases

Any deployed use case of the model - whether commercial or not - is currently out of scope. Non-deployed use cases such as image search in a constrained environment, are also not recommended unless there is thorough in-domain testing of the model with a specific, fixed class taxonomy. This is because our safety assessment demonstrated a high need for task specific testing especially given the variability of CLIP’s performance with different class taxonomies. This makes untested and unconstrained deployment of the model in any use case currently potentially harmful.

Certain use cases which would fall under the domain of surveillance and facial recognition are always out-of-scope regardless of performance of the model. This is because the use of artificial intelligence for tasks such as these can be premature currently given the lack of testing norms and checks to ensure its fair use.

Since the model has not been purposefully trained in or evaluated on any languages other than English, its use should be limited to English language use cases.

Data

The model was trained on publicly available image-caption data. This was done through a combination of crawling a handful of websites and using commonly-used pre-existing image datasets such as YFCC100M . A large portion of the data comes from our crawling of the internet. This means that the data is more representative of people and societies most connected to the internet which tend to skew towards more developed nations, and younger, male users.

Data Mission Statement

Our goal with building this dataset was to test out robustness and generalizability in computer vision tasks. As a result, the focus was on gathering large quantities of data from different publicly-available internet data sources. The data was gathered in a mostly non-interventionist manner. However, we only crawled websites that had policies against excessively violent and adult images and allowed us to filter out such content. We do not intend for this dataset to be used as the basis for any commercial or deployed model and will not be releasing the dataset.

Performance and Limitations
Performance

We have evaluated the performance of CLIP on a wide range of benchmarks across a variety of computer vision datasets such as OCR to texture recognition to fine-grained classification. The paper describes model performance on the following datasets:

  • Food101
  • CIFAR10
  • CIFAR100
  • Birdsnap
  • SUN397
  • Stanford Cars
  • FGVC Aircraft
  • VOC2007
  • DTD
  • Oxford-IIIT Pet dataset
  • Caltech101
  • Flowers102
  • MNIST
  • SVHN
  • IIIT5K
  • Hateful Memes
  • SST-2
  • UCF101
  • Kinetics700
  • Country211
  • CLEVR Counting
  • KITTI Distance
  • STL-10
  • RareAct
  • Flickr30
  • MSCOCO
  • ImageNet
  • ImageNet-A
  • ImageNet-R
  • ImageNet Sketch
  • ObjectNet (ImageNet Overlap)
  • Youtube-BB
  • ImageNet-Vid
Limitations

CLIP and our analysis of it have a number of limitations. CLIP currently struggles with respect to certain tasks such as fine grained classification and counting objects. CLIP also poses issues with regards to fairness and bias which we discuss in the paper and briefly in the next section. Additionally, our approach to testing CLIP also has an important limitation- in many cases we have used linear probes to evaluate the performance of CLIP and there is evidence suggesting that linear probes can underestimate model performance.

Bias and Fairness

We find that the performance of CLIP - and the specific biases it exhibits - can depend significantly on class design and the choices one makes for categories to include and exclude. We tested the risk of certain kinds of denigration with CLIP by classifying images of people from Fairface into crime-related and non-human animal categories. We found significant disparities with respect to race and gender. Additionally, we found that these disparities could shift based on how the classes were constructed. (Details captured in the Broader Impacts Section in the paper).

We also tested the performance of CLIP on gender, race and age classification using the Fairface dataset (We default to using race categories as they are constructed in the Fairface dataset.) in order to assess quality of performance across different demographics. We found accuracy >96% across all races for gender classification with ‘Middle Eastern’ having the highest accuracy (98.4%) and ‘White’ having the lowest (96.5%). Additionally, CLIP averaged ~93% for racial classification and ~63% for age classification. Our use of evaluations to test for gender, race and age classification as well as denigration harms is simply to evaluate performance of the model across people and surface potential risks and not to demonstrate an endorsement/enthusiasm for such tasks.

Pricing of clip-vit-large-patch14 replicate.com

Run time and cost

This model costs approximately $0.00022 to run on Replicate, or 4545 runs per $1, but this varies depending on your inputs. It is also open source and you can run it on your own computer with Docker .

This model runs on Nvidia T4 GPU hardware . Predictions typically complete within 1 seconds.

Runs of cjwbw clip-vit-large-patch14 on replicate.com

6.9M
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About clip-vit-large-patch14 replicate.com Model

clip-vit-large-patch14 replicate.com

clip-vit-large-patch14 replicate.com is an AI model on replicate.com that provides clip-vit-large-patch14's model effect (openai/clip-vit-large-patch14 with Transformers), which can be used instantly with this cjwbw clip-vit-large-patch14 model. replicate.com supports a free trial of the clip-vit-large-patch14 model, and also provides paid use of the clip-vit-large-patch14. Support call clip-vit-large-patch14 model through api, including Node.js, Python, http.

clip-vit-large-patch14 replicate.com Url

https://replicate.com/cjwbw/clip-vit-large-patch14

cjwbw clip-vit-large-patch14 online free

clip-vit-large-patch14 replicate.com is an online trial and call api platform, which integrates clip-vit-large-patch14's modeling effects, including api services, and provides a free online trial of clip-vit-large-patch14, you can try clip-vit-large-patch14 online for free by clicking the link below.

cjwbw clip-vit-large-patch14 online free url in replicate.com:

https://replicate.com/cjwbw/clip-vit-large-patch14

clip-vit-large-patch14 install

clip-vit-large-patch14 is an open source model from GitHub that offers a free installation service, and any user can find clip-vit-large-patch14 on GitHub to install. At the same time, replicate.com provides the effect of clip-vit-large-patch14 install, users can directly use clip-vit-large-patch14 installed effect in replicate.com for debugging and trial. It also supports api for free installation.

clip-vit-large-patch14 install url in replicate.com:

https://replicate.com/cjwbw/clip-vit-large-patch14

clip-vit-large-patch14 install url in github:

https://github.com/chenxwh/cog-clip

Url of clip-vit-large-patch14

clip-vit-large-patch14 replicate.com Url

clip-vit-large-patch14 Github

clip-vit-large-patch14 Owner Github

Provider of clip-vit-large-patch14 replicate.com

Other API from cjwbw

replicate

Remove images background

Total runs: 8.3M
Run Growth: 0
Growth Rate: 0.00%
Updated:November 30 2022
replicate

ZoeDepth: Combining relative and metric depth

Total runs: 4.5M
Run Growth: 0
Growth Rate: 0.00%
Updated:March 05 2023
replicate

Anime-themed text-to-image stable diffusion model

Total runs: 4.0M
Run Growth: 0
Growth Rate: 0.00%
Updated:March 20 2024
replicate

high-quality, highly detailed anime style stable-diffusion with better VAE

Total runs: 3.5M
Run Growth: 0
Growth Rate: 0.00%
Updated:January 15 2023
replicate

high-quality, highly detailed anime-style Stable Diffusion models

Total runs: 3.3M
Run Growth: 0
Growth Rate: 0.00%
Updated:January 23 2023
replicate

Real-ESRGAN: Real-World Blind Super-Resolution

Total runs: 2.2M
Run Growth: 0
Growth Rate: 0.00%
Updated:February 19 2023
replicate

powerful open-source visual language model

Total runs: 1.5M
Run Growth: 0
Growth Rate: 0.00%
Updated:November 30 2023
replicate

Dream Shaper stable diffusion

Total runs: 1.3M
Run Growth: 0
Growth Rate: 0.00%
Updated:March 12 2023
replicate

Stable Diffusion on Danbooru images

Total runs: 1.1M
Run Growth: 0
Growth Rate: 0.00%
Updated:October 10 2022
replicate

Colorization using a Generative Color Prior for Natural Images

Total runs: 564.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 04 2022
replicate

Real-ESRGAN super-resolution model from ruDALL-E

Total runs: 483.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:September 29 2022
replicate

Robust Monocular Depth Estimation

Total runs: 414.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:February 15 2023
replicate

high-quality, highly detailed anime style stable-diffusion

Total runs: 353.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:December 20 2022
replicate

sd-v2 with diffusers, test version!

Total runs: 280.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:December 02 2022
replicate

a dreambooth model trained on a diverse set of analog photographs

Total runs: 234.4K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 01 2023
replicate

Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild. This version uses LLaVA-13b for captioning.

Total runs: 186.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:February 24 2024
replicate

Demucs Music Source Separation

Total runs: 184.6K
Run Growth: 0
Growth Rate: 0.00%
Updated:July 02 2023
replicate

Advanced text-image comprehension and composition based on InternLM

Total runs: 164.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:October 02 2023
replicate

Multi-stage text-to-video generation

Total runs: 143.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:March 24 2023
replicate

Blind Face Restoration with Vector-Quantized Dictionary and Parallel Decoder

Total runs: 140.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 04 2022
replicate

Stylized Audio-Driven Single Image Talking Face Animation

Total runs: 127.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:June 01 2024
replicate

SeamlessM4T—Massively Multilingual & Multimodal Machine Translation

Total runs: 82.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:September 14 2023
replicate

VideoCrafter2: Text-to-Video and Image-to-Video Generation and Editing

Total runs: 66.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 31 2024
replicate

stable-diffusion with negative prompts, more scheduler

Total runs: 65.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:October 08 2022
replicate

Background removal model developed by BRIA.AI, trained on a carefully selected dataset and is available as an open-source model for non-commercial use.

Total runs: 55.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:February 07 2024
replicate

with large-v2 checkpoint

Total runs: 54.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:December 16 2022
replicate

Unsupervised Night Image Enhancement

Total runs: 41.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 14 2022
replicate

Text-to-Image Diffusion Models are Zero-Shot Video Generators

Total runs: 41.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 08 2023
replicate

stable-diffusion with v1-5 checkpoint

Total runs: 35.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:October 26 2022
replicate

Tuning-Free Multi-Subject Image Generation with Localized Attention

Total runs: 34.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:May 19 2023
replicate

high-quality highly detailed anime stylized latent diffusion model

Total runs: 31.8K
Run Growth: 0
Growth Rate: 0.00%
Updated:February 21 2023
replicate

mixed stable diffusion model

Total runs: 30.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:March 05 2023
replicate

Portraits with stable-diffusion

Total runs: 24.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 25 2023
replicate

VQ-Diffusion for Text-to-Image Synthesis

Total runs: 20.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 10 2022
replicate

Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild. This is the SUPIR-v0Q model and does NOT use LLaVA-13b.

Total runs: 19.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:February 24 2024
replicate

Image Manipulatinon with Diffusion Autoencoders

Total runs: 17.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 04 2022
replicate

Generating Conditional 3D Implicit Functions

Total runs: 15.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:May 20 2023
replicate

High-Quality Video Generation with Cascaded Latent Diffusion Models

Total runs: 13.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 21 2023
replicate

Audio-Driven Synthesis of Photorealistic Portrait Animations

Total runs: 13.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 01 2024
replicate

stable-diffusion models for high quality and detailed anime images

Total runs: 13.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 01 2023
replicate

Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild. This is the SUPIR-v0F model and does NOT use LLaVA-13b.

Total runs: 13.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:February 24 2024
replicate

Highly practical solution for robust monocular depth estimation by training on a combination of 1.5M labeled images and 62M+ unlabeled images

Total runs: 11.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 24 2024
replicate

Zero-Shot Speech Editing and Text-to-Speech in the Wild

Total runs: 9.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:March 15 2025
replicate

Pose-Invariant Hairstyle Transfer

Total runs: 9.6K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 21 2022
replicate

Point-E: A System for Generating 3D Point Clouds from Complex Prompts

Total runs: 8.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 17 2023
replicate

A linear estimator on top of clip to predict the aesthetic quality of pictures

Total runs: 8.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 18 2022
replicate

Decoding Micromotion in Low-dimensional Latent Spaces from StyleGAN

Total runs: 8.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 04 2022
replicate

fine-tuned Stable Diffusion model trained on the game art from Elden Ring

Total runs: 6.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 03 2022
replicate

Zero-shot Image-to-Image Translation

Total runs: 6.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:February 12 2023
replicate

Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Total runs: 6.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 14 2024
replicate

Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

Total runs: 6.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:March 29 2023
replicate

Van Gough on Stable Diffusion via Dreambooth

Total runs: 5.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 08 2022
replicate

Finte-tuned Stable Diffusion on high quality 3D images with a futuristic Sci-Fi theme

Total runs: 5.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 08 2023
replicate

face alignment using stylegan-encoding

Total runs: 4.8K
Run Growth: 0
Growth Rate: 0.00%
Updated:May 27 2022
replicate

Clip-Guided Diffusion Model for Image Generation

Total runs: 4.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:March 12 2022
replicate

Efficient Pretraining of Text-to-Image Models

Total runs: 4.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:September 16 2023
replicate

Separate Anything You Describe

Total runs: 4.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:October 20 2023
replicate

Inpainting using Denoising Diffusion Probabilistic Models

Total runs: 4.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 17 2022
replicate

Learning Adapters towards Controllable for Text-to-Image Diffusion Models

Total runs: 3.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:February 18 2023
replicate

End-to-End Document Image Enhancement Transformer

Total runs: 3.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 30 2022
replicate

dreambooth trained on a very diverse dataset ranging from photographs to paintings

Total runs: 3.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:December 09 2022
replicate

Disco Diffusion style on Stable Diffusion via Dreambooth

Total runs: 3.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 08 2022
replicate

Efficient Diffusion Model for Image Super-resolution by Residual Shifting

Total runs: 3.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 02 2023
replicate

Real-Time High-Resolution Background Matting

Total runs: 2.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 18 2022
replicate

Prompt-to-prompt image editing with cross-attention control

Total runs: 2.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:October 30 2022
replicate

Training-free Controllable Text-to-Video Generation

Total runs: 2.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:May 28 2023
replicate

lightweight text-to-speech (TTS) model, trained on 10.5K hours of audio data

Total runs: 2.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 16 2024
replicate

A Visual Language Model for GUI Agents

Total runs: 2.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:February 05 2024
replicate

herge_style on Stable Diffusion via Dreambooth

Total runs: 2.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 08 2022
replicate

Controlling Vision-Language Models for Universal Image Restoration

Total runs: 2.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:October 13 2023
replicate

Consistent Diffusion Features for Consistent Video Editing

Total runs: 2.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:January 23 2024
replicate

Finetuned Stable-diffusion from Gerry Anderson Supermarionation

Total runs: 1.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:March 04 2023
replicate

text-to-image generation

Total runs: 1.8K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 10 2022
replicate

Diffusion Models as Text Painters

Total runs: 1.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:June 04 2023
replicate

Open-source Distilled Stable Diffusion 100% speedup

Total runs: 1.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:December 13 2023
replicate

Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis

Total runs: 1.4K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 27 2024
replicate

High-quality multilingual text-to-speech library

Total runs: 1.4K
Run Growth: 0
Growth Rate: 0.00%
Updated:March 03 2024
replicate

Panoptic Scene Graph Generation

Total runs: 1.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 13 2022
replicate

text-to-video generation model

Total runs: 1.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 26 2023