keras / clip_vit_base_patch32

huggingface.co
Total runs: 62
24-hour runs: 0
7-day runs: -62
30-day runs: -55
Model's Last Updated: October 28 2025

Introduction of clip_vit_base_patch32

Model Details of clip_vit_base_patch32

Model Overview

Model Summary

This model is a CLIP (Contrastive Language-Image Pre-training) neural network. CLIP revolutionizes image understanding by learning visual concepts from natural language descriptions found online. It's been trained on a massive dataset of image-text pairs, allowing it to excel at tasks like zero-shot image classification, image search based on text queries, and robust visual understanding. With CLIP, you can explore the power of aligning image and text representations within a shared embedding space.

Weights are released under the MIT License . Keras model code is released under the Apache 2 License .

Links
Installation

Keras and KerasCV can be installed with:

pip install -U -q keras-cv
pip install -U -q keras>=3

Jax, TensorFlow, and Torch come preinstalled in Kaggle Notebooks. For instructions on installing them in another environment see the Keras Getting Started page.

Presets

The following model checkpoints are provided by the Keras team. Full code examples for each are available below.

Preset name Parameters Description
clip-vit-base-patch16 149.62M The model uses a ViT-B/16 Transformer architecture as an image encoder and uses a masked self-attention Transformer as a text encoder. These encoders are trained to maximize the similarity of (image, text) pairs via a contrastive loss. The model uses a patch size of 16 and input images of size (224, 224)
clip-vit-base-patch32 151.28M The model uses a ViT-B/32 Transformer architecture as an image encoder and uses a masked self-attention Transformer as a text encoder. These encoders are trained to maximize the similarity of (image, text) pairs via a contrastive loss.The model uses a patch size of 32 and input images of size (224, 224)
clip-vit-large-patch14 427.62M The model uses a ViT-L/14 Transformer architecture as an image encoder and uses a masked self-attention Transformer as a text encoder. These encoders are trained to maximize the similarity of (image, text) pairs via a contrastive loss.The model uses a patch size of 14 and input images of size (224, 224)
clip-vit-large-patch14-336 427.94M The model uses a ViT-L/14 Transformer architecture as an image encoder and uses a masked self-attention Transformer as a text encoder. These encoders are trained to maximize the similarity of (image, text) pairs via a contrastive loss.The model uses a patch size of 14 and input images of size (336, 336)
Example code
from keras import ops
import keras
from keras_cv.models.feature_extractor.clip import CLIPProcessor
from keras_cv.models import CLIP

processor = CLIPProcessor("vocab.json", "merges.txt")
# processed_image = transform_image("cat.jpg", 224)
tokens = processor(["mountains", "cat on tortoise", "house"])
model = CLIP.from_preset("clip-vit-base-patch32")
output = model({
                "images": processed_image,
                "token_ids": tokens['token_ids'],
                "padding_mask": tokens['padding_mask']})


# optional if you need to pre process image
def transform_image(image_path, input_resolution):
    mean = ops.array([0.48145466, 0.4578275, 0.40821073])
    std = ops.array([0.26862954, 0.26130258, 0.27577711])

    image = keras.utils.load_img(image_path)
    image = keras.utils.img_to_array(image)
    image = (
        ops.image.resize(
            image,
            (input_resolution, input_resolution),
            interpolation="bicubic",
        )
        / 255.0
    )
    central_fraction = input_resolution / image.shape[0]
    width, height = image.shape[0], image.shape[1]
    left = ops.cast((width - width * central_fraction) / 2, dtype="int32")
    top = ops.cast((height - height * central_fraction) / 2, dtype="int32")
    right = ops.cast((width + width * central_fraction) / 2, dtype="int32")
    bottom = ops.cast(
        (height + height * central_fraction) / 2, dtype="int32"
    )

    image = ops.slice(
        image, [left, top, 0], [right - left, bottom - top, 3]
    )

    image = (image - mean) / std
    return ops.expand_dims(image, axis=0)

Runs of keras clip_vit_base_patch32 on huggingface.co

62
Total runs
0
24-hour runs
-30
3-day runs
-62
7-day runs
-55
30-day runs

More Information About clip_vit_base_patch32 huggingface.co Model

clip_vit_base_patch32 huggingface.co

clip_vit_base_patch32 huggingface.co is an AI model on huggingface.co that provides clip_vit_base_patch32's model effect (), which can be used instantly with this keras clip_vit_base_patch32 model. huggingface.co supports a free trial of the clip_vit_base_patch32 model, and also provides paid use of the clip_vit_base_patch32. Support call clip_vit_base_patch32 model through api, including Node.js, Python, http.

clip_vit_base_patch32 huggingface.co Url

https://huggingface.co/keras/clip_vit_base_patch32

keras clip_vit_base_patch32 online free

clip_vit_base_patch32 huggingface.co is an online trial and call api platform, which integrates clip_vit_base_patch32's modeling effects, including api services, and provides a free online trial of clip_vit_base_patch32, you can try clip_vit_base_patch32 online for free by clicking the link below.

keras clip_vit_base_patch32 online free url in huggingface.co:

https://huggingface.co/keras/clip_vit_base_patch32

clip_vit_base_patch32 install

clip_vit_base_patch32 is an open source model from GitHub that offers a free installation service, and any user can find clip_vit_base_patch32 on GitHub to install. At the same time, huggingface.co provides the effect of clip_vit_base_patch32 install, users can directly use clip_vit_base_patch32 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

clip_vit_base_patch32 install url in huggingface.co:

https://huggingface.co/keras/clip_vit_base_patch32

Url of clip_vit_base_patch32

clip_vit_base_patch32 huggingface.co Url

Provider of clip_vit_base_patch32 huggingface.co

keras
ORGANIZATIONS

Other API from keras

huggingface.co

Total runs: 22
Run Growth: 19
Growth Rate: 86.36%
Updated:February 27 2026
huggingface.co

Total runs: 22
Run Growth: 12
Growth Rate: 54.55%
Updated:March 25 2025
huggingface.co

Total runs: 13
Run Growth: -28
Growth Rate: -215.38%
Updated:June 17 2025
huggingface.co

Total runs: 13
Run Growth: 6
Growth Rate: 46.15%
Updated:February 27 2026
huggingface.co

Total runs: 13
Run Growth: 10
Growth Rate: 76.92%
Updated:March 25 2025
huggingface.co

Total runs: 13
Run Growth: -23
Growth Rate: -176.92%
Updated:June 17 2025
huggingface.co

Total runs: 11
Run Growth: 8
Growth Rate: 72.73%
Updated:February 27 2026
huggingface.co

Total runs: 9
Run Growth: 3
Growth Rate: 33.33%
Updated:March 25 2025
huggingface.co

Total runs: 8
Run Growth: 1
Growth Rate: 12.50%
Updated:March 25 2025
huggingface.co

Total runs: 7
Run Growth: -12
Growth Rate: -171.43%
Updated:May 15 2026
huggingface.co

Total runs: 7
Run Growth: -27
Growth Rate: -385.71%
Updated:June 17 2025
huggingface.co

Total runs: 7
Run Growth: 4
Growth Rate: 57.14%
Updated:February 27 2026