This model is a fine-tuned version of
google/gemma-3n-E2B-it
adapted for the Bambara (Bamanankan) language, one of the major languages spoken in Mali and surrounding West African countries. It was trained on the
sudoping01/bambara-instructions
dataset.
Model Description
MALIBA-LLM is a LoRA-adapted Gemma 3 model designed to provide instruction-following capabilities in Bambara, a language spoken by over 14 million people across Mali, Côte d'Ivoire, Burkina Faso, and other West African countries. This model represents a significant step in democratizing AI access for speakers of this Niger-Congo language.
Key Features
Language Support
: Primary support for Bambara (Bamanankan) with capability for code-switching with French when appropriate for technical terms
Model Type
: LoRA adapter for Gemma 3 (2.7B instruction-tuned)
Training Method
: Supervised fine-tuning using a cleaned subset of the MALIBA-Instructions dataset
Context Length
: 4096 tokens
Capabilities
: Conversational assistance, instruction following, knowledge retrieval, creative content generation, and reasoning in Bambara
Intended Uses & Limitations
Intended Uses
Conversational assistance in Bambara
Content generation for educational materials in Bambara
Translation assistance between Bambara and other languages
Cultural context preservation for Bambara-speaking regions
Improving digital accessibility for Bambara speakers
Limitations
This is an experimental model and the first open-source effort for Bambara LLM development
Performance varies across different task types, with stronger performance in conversational and general knowledge domains
Limited technical vocabulary in pure Bambara (uses French code-switching for some technical terms)
The model may occasionally produce grammatical structures that deviate from standard Bambara, particularly for complex sentences
No built-in safety guardrails specific to Bambara cultural context (inherits base model safety mechanisms)
Training and Evaluation Data
Training Data
The model was trained on the
sudoping01/bambara-instructions
dataset, specifically using the cleaned subset which contains 563,892 high-quality conversations. This dataset was created using a novel methodology that combines linguistic knowledge with LLM reasoning capabilities to transform diverse instruction datasets into Bambara while preserving grammatical structures and cultural context.
The original datasets transformed for this training include:
The model was evaluated on a held-out validation set (1% of the training data) and achieved a final validation loss of 0.4952. Additional human evaluation by native Bambara speakers verified the quality of outputs across various domains.
Key evaluation metrics:
Loss on validation set: 0.4952
Memory usage during training: 57.85 GiB
Training Procedure
Training Hyperparameters
The model was trained using the following hyperparameters:
Learning rate: 0.00012
Train batch size: 8 (per device)
Total train batch size: 128 (with gradient accumulation and distributed training)
Optimizer: AdamW (8-bit)
LR scheduler: Cosine with warmup
Training steps: 21,126
Epochs: 3
LoRA rank (r): 64
LoRA alpha: 128
LoRA dropout: 0.05
Sequence length: 4096
Training Results
Training Loss
Epoch
Step
Validation Loss
Mem Active(GiB)
No log
0
0
7.4595
19.86
0.8265
0.5
3521
0.7787
57.85
0.7107
1.0
7042
0.6745
57.85
0.6363
1.5
10563
0.6026
57.85
0.5421
2.0
14084
0.5429
57.85
0.5733
2.5
17605
0.5039
57.85
0.5401
3.0
21126
0.4952
57.85
The training demonstrated consistent improvement across epochs with validation loss decreasing from 7.4595 to 0.4952.
Usage
Loading the Model
from peft import PeftModel, PeftConfig
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load base model and adapter
model_name = "sudoping01/bambara-llm-exp3"
config = PeftConfig.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
config.base_model_name_or_path,
torch_dtype=torch.bfloat16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(config.base_model_name_or_path)
model = PeftModel.from_pretrained(model, model_name)
Sample Conversation
# Using the Gemma3 chat template
messages = [
{"role": "user", "content": "I ni ce! I bɛ di? I bɛ se ka n dɛmɛ ka Bamanankan kalan wa?"}
]
# Format prompt with chat template
prompt = tokenizer.apply_chat_template(messages, tokenize=False)
# Generate response
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.7,
top_p=0.9,
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
Ethical Considerations
This model aims to increase accessibility of AI technology to Bambara speakers while respecting cultural context. However, as with all language models, care should be taken regarding:
Cultural Sensitivity
: The model was trained on transformed data, which may not fully represent all Bambara cultural contexts and regional variations.
Content Generation Risks
: The model may generate content that could be considered inappropriate in specific cultural contexts. Users should be aware of local cultural norms.
Digital Divide
: While this model helps bridge linguistic divides, access to the technology required to use this model may still be limited in some Bambara-speaking communities.
Dialectal Variation
: Bambara has several dialects, and this model may not perform equally well across all regional variations.
Additional Information
This model was developed as part of the MALIBA project, which aims to increase representation of African languages in AI systems. The complete implementation pipeline is available at
https://github.com/sudoping01/instructions-gen
.
The work is still in progress, with ongoing improvements planned for future releases.
Citation
@article{diallo2025bambara,
title={Linguistically-Informed Large Language Models for Low-Resource Instruction Dataset Creation},
author={Diallo, Seydou},
journal={Unpublished manuscript},
year={2025},
month={July}
}
Framework Versions
PEFT 0.17.0
Transformers 4.55.2
PyTorch 2.6.0+cu124
Datasets 4.0.0
Tokenizers 0.21.4
Runs of sudoping01 maliba-llm on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About maliba-llm huggingface.co Model
maliba-llm huggingface.co is an AI model on huggingface.co that provides maliba-llm's model effect (), which can be used instantly with this sudoping01 maliba-llm model. huggingface.co supports a free trial of the maliba-llm model, and also provides paid use of the maliba-llm. Support call maliba-llm model through api, including Node.js, Python, http.
maliba-llm huggingface.co is an online trial and call api platform, which integrates maliba-llm's modeling effects, including api services, and provides a free online trial of maliba-llm, you can try maliba-llm online for free by clicking the link below.
sudoping01 maliba-llm online free url in huggingface.co:
maliba-llm is an open source model from GitHub that offers a free installation service, and any user can find maliba-llm on GitHub to install. At the same time, huggingface.co provides the effect of maliba-llm install, users can directly use maliba-llm installed effect in huggingface.co for debugging and trial. It also supports api for free installation.