🌟 This is a small (only 1.2B parameters) visual language model on Hugging Face that responds to Traditional Chinese instructions given an image input! 🌟
✨ Developed compatible with the Transformers library, TaiVisionLM is quick to load, fine-tune, and use for lightning-fast inferences without needing any external libraries! ⚡️
Ready to experience the Traditional Chinese visual language model? Let's go! 🖼️🤖
This model is a multimodal large language model that combines
SigLIP
as its vision encoder with
Tinyllama
as its language model. The vision projector connects the two modalities together.
Its architecture closely resembles
PaliGemma
.
We trained the vision projector and language model using LoRA using 1M image-text pairs to align visual and textual features.
This model is the finetuned version of
benchang1110/TaiVisionLM-base-v1
. We fintuned the model using 1M image-text pairs. The finetuned model will generate a longer and more detailed description of the image.
Task Specific Training
The aligned model undergoes further training for tasks such as short captioning, detailed captioning, and simple visual question answering.
We will undergo this stage after the dataset is ready!
In Transformers, you can load the model and do inference as follows:
IMPORTANT NOTE:
TaiVisionLM model is not yet integrated natively into the Transformers library. So you need to set
trust_remote_code=True
when loading the model. It will download the
configuration_taivisionlm.py
,
modeling_taivisionlm.py
and
processing_taivisionlm.py
files from the repo. You can check out the content of these files under the
Files and Versions
tab and pin the specific versions if you have any concerns regarding malicious code.
Since we don't have enough resources to train the model on the whole dataset, we only use 250k image-text pairs for training. The following training hyperparameters are used in feature alignment and task specific training stages respectively:
Feature Alignment
Data size
Global Batch Size
Learning Rate
Epochs
Max Length
Weight Decay
250k
2
5e-5
1
2048
1e-5
We use full-parameter finetuning for the projector and apply LoRA to the language model.
We will update the training procedure once we have more resources to train the model on the whole dataset.
Compute Infrastructure
Feature Alignment
1xV100(32GB), took approximately 12 GPU hours.
Runs of benchang1110 TaiVisionLM-base-v2 on huggingface.co
11
Total runs
0
24-hour runs
3
3-day runs
1
7-day runs
-3
30-day runs
More Information About TaiVisionLM-base-v2 huggingface.co Model
TaiVisionLM-base-v2 huggingface.co
TaiVisionLM-base-v2 huggingface.co is an AI model on huggingface.co that provides TaiVisionLM-base-v2's model effect (), which can be used instantly with this benchang1110 TaiVisionLM-base-v2 model. huggingface.co supports a free trial of the TaiVisionLM-base-v2 model, and also provides paid use of the TaiVisionLM-base-v2. Support call TaiVisionLM-base-v2 model through api, including Node.js, Python, http.
TaiVisionLM-base-v2 huggingface.co is an online trial and call api platform, which integrates TaiVisionLM-base-v2's modeling effects, including api services, and provides a free online trial of TaiVisionLM-base-v2, you can try TaiVisionLM-base-v2 online for free by clicking the link below.
benchang1110 TaiVisionLM-base-v2 online free url in huggingface.co:
TaiVisionLM-base-v2 is an open source model from GitHub that offers a free installation service, and any user can find TaiVisionLM-base-v2 on GitHub to install. At the same time, huggingface.co provides the effect of TaiVisionLM-base-v2 install, users can directly use TaiVisionLM-base-v2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
TaiVisionLM-base-v2 install url in huggingface.co: