MobileVLM is a competent multimodal vision language model (MMVLM) targeted to run on mobile devices. It is an amalgamation of a myriad of architectural designs and techniques that are mobile-oriented, which comprises a set of language models at the scale of 1.4B and 2.7B parameters, trained from scratch, a multimodal vision model that is pre-trained in the CLIP fashion, cross-modality interaction via an efficient projector. We evaluate MobileVLM on several typical VLM benchmarks. Our models demonstrate on par performance compared with a few much larger models. More importantly, we measure the inference speed on both a Qualcomm Snapdragon 888 CPU and an NVIDIA Jeston Orin GPU, and we obtain state-of-the-art performance of 21.5 tokens and 65.3 tokens per second, respectively.
The MobileVLM-3B was built on our
MobileLLaMA-2.7B-Chat
to facilitate the off-the-shelf deployment.
MobileVLM-3B huggingface.co is an AI model on huggingface.co that provides MobileVLM-3B's model effect (), which can be used instantly with this mtgv MobileVLM-3B model. huggingface.co supports a free trial of the MobileVLM-3B model, and also provides paid use of the MobileVLM-3B. Support call MobileVLM-3B model through api, including Node.js, Python, http.
MobileVLM-3B huggingface.co is an online trial and call api platform, which integrates MobileVLM-3B's modeling effects, including api services, and provides a free online trial of MobileVLM-3B, you can try MobileVLM-3B online for free by clicking the link below.
mtgv MobileVLM-3B online free url in huggingface.co:
MobileVLM-3B is an open source model from GitHub that offers a free installation service, and any user can find MobileVLM-3B on GitHub to install. At the same time, huggingface.co provides the effect of MobileVLM-3B install, users can directly use MobileVLM-3B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.