Aya 23 is an open weights research release of an instruction fine-tuned model with highly advanced multilingual capabilities. Aya 23 focuses on pairing a highly performant pre-trained
Command family
of models with the recently released
Aya Collection
. The result is a powerful multilingual large language model serving 23 languages.
This model card corresponds to the 8-billion version of the Aya 23 model. We also released a 35-billion version which you can find
here
.
You can try out Aya 23 (35B) before downloading the weights in our hosted Hugging Face Space
here
.
Usage
Please install transformers from the source repository that includes the necessary changes for this model
# pip install transformers==4.41.1from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "CohereForAI/aya-23-8B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
# Format message with the command-r-plus chat template
messages = [{"role": "user", "content": "Anneme onu ne kadar sevdiğimi anlatan bir mektup yaz"}]
input_ids = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt")
## <BOS_TOKEN><|START_OF_TURN_TOKEN|><|USER_TOKEN|>Anneme onu ne kadar sevdiğimi anlatan bir mektup yaz<|END_OF_TURN_TOKEN|><|START_OF_TURN_TOKEN|><|CHATBOT_TOKEN|>
gen_tokens = model.generate(
input_ids,
max_new_tokens=100,
do_sample=True,
temperature=0.3,
)
gen_text = tokenizer.decode(gen_tokens[0])
print(gen_text)
Example Notebook
This notebook
showcases a detailed use of Aya 23 (8B) including inference and fine-tuning with
QLoRA
.
Model Details
Input
: Models input text only.
Output
: Models generate text only.
Model Architecture
: Aya-23-8B is an auto-regressive language model that uses an optimized transformer architecture. After pretraining, this model is fine-tuned (IFT) to follow human instructions.
Languages covered
: The model is particularly optimized for multilinguality and supports the following languages: Arabic, Chinese (simplified & traditional), Czech, Dutch, English, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Korean, Persian, Polish, Portuguese, Romanian, Russian, Spanish, Turkish, Ukrainian, and Vietnamese
Context length
: 8192
Evaluation
Please refer to the
Aya 23 technical report
for further details about the base model, data, instruction tuning, and evaluation.
Terms of Use
We hope that the release of this model will make community-based research efforts more accessible, by releasing the weights of a highly performant multilingual model to researchers all over the world. This model is governed by a
CC-BY-NC
License with an acceptable use addendum, and also requires adhering to
C4AI's Acceptable Use Policy
.
Runs of QuantFactory aya-23-8B-GGUF on huggingface.co
913
Total runs
10
24-hour runs
99
3-day runs
221
7-day runs
754
30-day runs
More Information About aya-23-8B-GGUF huggingface.co Model
aya-23-8B-GGUF huggingface.co is an AI model on huggingface.co that provides aya-23-8B-GGUF's model effect (), which can be used instantly with this QuantFactory aya-23-8B-GGUF model. huggingface.co supports a free trial of the aya-23-8B-GGUF model, and also provides paid use of the aya-23-8B-GGUF. Support call aya-23-8B-GGUF model through api, including Node.js, Python, http.
aya-23-8B-GGUF huggingface.co is an online trial and call api platform, which integrates aya-23-8B-GGUF's modeling effects, including api services, and provides a free online trial of aya-23-8B-GGUF, you can try aya-23-8B-GGUF online for free by clicking the link below.
QuantFactory aya-23-8B-GGUF online free url in huggingface.co:
aya-23-8B-GGUF is an open source model from GitHub that offers a free installation service, and any user can find aya-23-8B-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of aya-23-8B-GGUF install, users can directly use aya-23-8B-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.