Disclaimer: The team releasing MobileViT did not write a model card for this model so this model card has been written by the Hugging Face team.
Model description
MobileViT is a light-weight, low latency convolutional neural network that combines MobileNetV2-style layers with a new block that replaces local processing in convolutions with global processing using transformers. As with ViT (Vision Transformer), the image data is converted into flattened patches before it is processed by the transformer layers. Afterwards, the patches are "unflattened" back into feature maps. This allows the MobileViT-block to be placed anywhere inside a CNN. MobileViT does not require any positional embeddings.
The model in this repo adds a
DeepLabV3
head to the MobileViT backbone for semantic segmentation.
Intended uses & limitations
You can use the raw model for semantic segmentation. See the
model hub
to look for fine-tuned versions on a task that interests you.
Currently, both the feature extractor and model support PyTorch.
Training data
The MobileViT + DeepLabV3 model was pretrained on
ImageNet-1k
, a dataset consisting of 1 million images and 1,000 classes, and then fine-tuned on the
PASCAL VOC2012
dataset.
Training procedure
Preprocessing
At inference time, images are center-cropped at 512x512. Pixels are normalized to the range [0, 1]. Images are expected to be in BGR pixel order, not RGB.
Pretraining
The MobileViT networks are trained from scratch for 300 epochs on ImageNet-1k on 8 NVIDIA GPUs with an effective batch size of 1024 and learning rate warmup for 3k steps, followed by cosine annealing. Also used were label smoothing cross-entropy loss and L2 weight decay. Training resolution varies from 160x160 to 320x320, using multi-scale sampling.
To obtain the DeepLabV3 model, MobileViT was fine-tuned on the PASCAL VOC dataset using 4 NVIDIA A100 GPUs.
@inproceedings{vision-transformer,
title = {MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer},
author = {Sachin Mehta and Mohammad Rastegari},
year = {2022},
URL = {https://arxiv.org/abs/2110.02178}
}
Runs of apple deeplabv3-mobilevit-x-small on huggingface.co
212
Total runs
-5
24-hour runs
1
3-day runs
36
7-day runs
115
30-day runs
More Information About deeplabv3-mobilevit-x-small huggingface.co Model
More deeplabv3-mobilevit-x-small license Visit here:
deeplabv3-mobilevit-x-small huggingface.co is an AI model on huggingface.co that provides deeplabv3-mobilevit-x-small's model effect (), which can be used instantly with this apple deeplabv3-mobilevit-x-small model. huggingface.co supports a free trial of the deeplabv3-mobilevit-x-small model, and also provides paid use of the deeplabv3-mobilevit-x-small. Support call deeplabv3-mobilevit-x-small model through api, including Node.js, Python, http.
deeplabv3-mobilevit-x-small huggingface.co is an online trial and call api platform, which integrates deeplabv3-mobilevit-x-small's modeling effects, including api services, and provides a free online trial of deeplabv3-mobilevit-x-small, you can try deeplabv3-mobilevit-x-small online for free by clicking the link below.
apple deeplabv3-mobilevit-x-small online free url in huggingface.co:
deeplabv3-mobilevit-x-small is an open source model from GitHub that offers a free installation service, and any user can find deeplabv3-mobilevit-x-small on GitHub to install. At the same time, huggingface.co provides the effect of deeplabv3-mobilevit-x-small install, users can directly use deeplabv3-mobilevit-x-small installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
deeplabv3-mobilevit-x-small install url in huggingface.co: