C4AI Aya Vision 8B
is an open weights research release of an 8-billion parameter model with advanced capabilities optimized for a variety of vision-language use cases, including OCR, captioning, visual reasoning, summarization, question answering, code, and more.
It is a multilingual model trained to excel in 23 languages in vision and language.
This model card corresponds to the 8-billion version of the Aya Vision model. We also released a 32-billion version which you can find
here
.
You can also talk to Aya Vision through the popular messaging service WhatsApp. Use this
link
to open a WhatsApp chatbox with Aya Vision.
If you don’t have WhatsApp downloaded on your machine you might need to do that, or, if you have it on your phone, you can follow the on-screen instructions to link your phone and WhatsApp Web.
By the end, you should see a text window which you can use to chat with the model.
More details about our WhatsApp integration are available
here
.
Example Notebook
You can also check out the following
notebook
to understand how to use Aya Vision for different use cases.
How to Use Aya Vision
Please install
transformers
from the source repository that includes the necessary changes for this model:
You can also use the model directly using transformers
pipeline
abstraction:
from transformers import pipeline
pipe = pipeline(model="CohereForAI/aya-vision-8b", task="image-text-to-text", device_map="auto")
# Format message with the aya-vision chat template
messages = [
{"role": "user",
"content": [
{"type": "image", "url": "https://media.istockphoto.com/id/458012057/photo/istanbul-turkey.jpg?s=612x612&w=0&k=20&c=qogAOVvkpfUyqLUMr_XJQyq-HkACXyYUSZbKhBlPrxo="},
{"type": "text", "text": "Bu resimde hangi anıt gösterilmektedir?"},
]},
]
outputs = pipe(text=messages, max_new_tokens=300, return_full_text=False)
print(outputs)
Model Details
Input:
Model accepts input text and images.
Output:
Model generates text.
Model Architecture:
This is a vision-language model that uses a multilingual language model based on
C4AI Command R7B
and further post-trained with the
Aya Expanse recipe
, paired with
SigLIP2-patch14-384
vision encoder through a multimodal adapter for vision-language understanding.
Image Processing:
We use
169 visual tokens
to encode an image tile with a resolution of
364x364 pixels
. Input images of arbitrary sizes are mapped to the nearest supported resolution based on the aspect ratio. Aya Vision uses up to 12 input tiles and a thumbnail (resized to 364x364) (2197 image tokens).
Languages covered:
The model has been trained on 23 languages: English, French, Spanish, Italian, German, Portuguese, Japanese, Korean, Arabic, Chinese (Simplified and Traditional), Russian, Polish, Turkish, Vietnamese, Dutch, Czech, Indonesian, Ukrainian, Romanian, Greek, Hindi, Hebrew, and Persian.
Context length
: Aya Vision 8B supports a context length of 16K.
For more details about how the model was trained, check out
our blogpost
.
We also evaluated Aya Vision 8B’s performance for text-only input against the same models using
m-ArenaHard
, a challenging open-ended generation evaluation, measured using win-rates using gpt-4o-2024-11-20 as a judge.
Model Card Contact
For errors or additional questions about details in this model card, contact
[email protected]
.
Terms of Use
We hope that the release of this model will make community-based research efforts more accessible by releasing the weights of a highly performant 8 billion parameter Vision-Language Model to researchers all over the world.
aya-vision-8b huggingface.co is an AI model on huggingface.co that provides aya-vision-8b's model effect (), which can be used instantly with this CohereLabs aya-vision-8b model. huggingface.co supports a free trial of the aya-vision-8b model, and also provides paid use of the aya-vision-8b. Support call aya-vision-8b model through api, including Node.js, Python, http.
aya-vision-8b huggingface.co is an online trial and call api platform, which integrates aya-vision-8b's modeling effects, including api services, and provides a free online trial of aya-vision-8b, you can try aya-vision-8b online for free by clicking the link below.
CohereLabs aya-vision-8b online free url in huggingface.co:
aya-vision-8b is an open source model from GitHub that offers a free installation service, and any user can find aya-vision-8b on GitHub to install. At the same time, huggingface.co provides the effect of aya-vision-8b install, users can directly use aya-vision-8b installed effect in huggingface.co for debugging and trial. It also supports api for free installation.