sudoping01 / maliba-llm

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: November 10 2025
text-generation

Introduction of maliba-llm

Model Details of maliba-llm

Built with Axolotl

See axolotl config

axolotl version: 0.12.2

base_model: google/gemma-3n-E2B-it
hub_model_id: sudoping01/bambara-llm-exp3 
plugins:
  - axolotl.integrations.cut_cross_entropy.CutCrossEntropyPlugin
cut_cross_entropy: true
load_in_4bit: false  
gradient_checkpointing: true
gradient_checkpointing_kwargs:
  use_reentrant: false
ddp: true
chat_template: gemma3n
eot_tokens:
  - <end_of_turn>
special_tokens:
  eot_token: <end_of_turn>
datasets:
  - path: sudoping01/bambara-instructions
    type: chat_template
    split: train
    name: cleaned
    field_messages: messages
    message_property_mappings:
      role: role
      content: content
val_set_size: 0.01
output_dir: ./outputs/bambara-gemma3n-lora-exp4
adapter: lora  
lora_r: 64     
lora_alpha: 128 
lora_dropout: 0.05
lora_target_modules: 'model.language_model.layers.[\d]+.(mlp|self_attn).(up|down|gate|q|k|v|o)_proj'
sequence_len: 4096 
sample_packing: false
pad_to_sequence_len: false
micro_batch_size: 8  
gradient_accumulation_steps: 2
num_epochs: 3  
optimizer: adamw_8bit
lr_scheduler: cosine
learning_rate: 1.2e-4  
warmup_ratio: 0.03
weight_decay: 0.01
bf16: auto
tf32: false
logging_steps: 10
saves_per_epoch: 2  
evals_per_epoch: 2

MALIBA-LLM: Bambara-LLM-Exp3

This model is a fine-tuned version of google/gemma-3n-E2B-it adapted for the Bambara (Bamanankan) language, one of the major languages spoken in Mali and surrounding West African countries. It was trained on the sudoping01/bambara-instructions dataset.

Model Description

MALIBA-LLM is a LoRA-adapted Gemma 3 model designed to provide instruction-following capabilities in Bambara, a language spoken by over 14 million people across Mali, Côte d'Ivoire, Burkina Faso, and other West African countries. This model represents a significant step in democratizing AI access for speakers of this Niger-Congo language.

Key Features
  • Language Support : Primary support for Bambara (Bamanankan) with capability for code-switching with French when appropriate for technical terms
  • Model Type : LoRA adapter for Gemma 3 (2.7B instruction-tuned)
  • Training Method : Supervised fine-tuning using a cleaned subset of the MALIBA-Instructions dataset
  • Context Length : 4096 tokens
  • Capabilities : Conversational assistance, instruction following, knowledge retrieval, creative content generation, and reasoning in Bambara
Intended Uses & Limitations
Intended Uses
  • Conversational assistance in Bambara
  • Content generation for educational materials in Bambara
  • Translation assistance between Bambara and other languages
  • Cultural context preservation for Bambara-speaking regions
  • Improving digital accessibility for Bambara speakers
Limitations
  • This is an experimental model and the first open-source effort for Bambara LLM development
  • Performance varies across different task types, with stronger performance in conversational and general knowledge domains
  • Limited technical vocabulary in pure Bambara (uses French code-switching for some technical terms)
  • The model may occasionally produce grammatical structures that deviate from standard Bambara, particularly for complex sentences
  • No built-in safety guardrails specific to Bambara cultural context (inherits base model safety mechanisms)
Training and Evaluation Data
Training Data

The model was trained on the sudoping01/bambara-instructions dataset, specifically using the cleaned subset which contains 563,892 high-quality conversations. This dataset was created using a novel methodology that combines linguistic knowledge with LLM reasoning capabilities to transform diverse instruction datasets into Bambara while preserving grammatical structures and cultural context.

The original datasets transformed for this training include:

  • Anthropic's RLHF datasets
  • OpenAssistant conversational data
  • Google's SMOL low-resource language dataset
  • Specialized instruction datasets (mathematical reasoning, code generation)
  • French-language instruction datasets
Evaluation

The model was evaluated on a held-out validation set (1% of the training data) and achieved a final validation loss of 0.4952. Additional human evaluation by native Bambara speakers verified the quality of outputs across various domains.

Key evaluation metrics:

  • Loss on validation set: 0.4952
  • Memory usage during training: 57.85 GiB
Training Procedure
Training Hyperparameters

The model was trained using the following hyperparameters:

  • Learning rate: 0.00012
  • Train batch size: 8 (per device)
  • Total train batch size: 128 (with gradient accumulation and distributed training)
  • Optimizer: AdamW (8-bit)
  • LR scheduler: Cosine with warmup
  • Training steps: 21,126
  • Epochs: 3
  • LoRA rank (r): 64
  • LoRA alpha: 128
  • LoRA dropout: 0.05
  • Sequence length: 4096
Training Results
Training Loss Epoch Step Validation Loss Mem Active(GiB)
No log 0 0 7.4595 19.86
0.8265 0.5 3521 0.7787 57.85
0.7107 1.0 7042 0.6745 57.85
0.6363 1.5 10563 0.6026 57.85
0.5421 2.0 14084 0.5429 57.85
0.5733 2.5 17605 0.5039 57.85
0.5401 3.0 21126 0.4952 57.85

The training demonstrated consistent improvement across epochs with validation loss decreasing from 7.4595 to 0.4952.

Usage
Loading the Model
from peft import PeftModel, PeftConfig
from transformers import AutoModelForCausalLM, AutoTokenizer

# Load base model and adapter
model_name = "sudoping01/bambara-llm-exp3"
config = PeftConfig.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    config.base_model_name_or_path,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(config.base_model_name_or_path)
model = PeftModel.from_pretrained(model, model_name)
Sample Conversation
# Using the Gemma3 chat template
messages = [
    {"role": "user", "content": "I ni ce! I bɛ di? I bɛ se ka n dɛmɛ ka Bamanankan kalan wa?"}
]

# Format prompt with chat template
prompt = tokenizer.apply_chat_template(messages, tokenize=False)

# Generate response
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
    **inputs,
    max_new_tokens=512,
    temperature=0.7,
    top_p=0.9,
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
Ethical Considerations

This model aims to increase accessibility of AI technology to Bambara speakers while respecting cultural context. However, as with all language models, care should be taken regarding:

  1. Cultural Sensitivity : The model was trained on transformed data, which may not fully represent all Bambara cultural contexts and regional variations.

  2. Content Generation Risks : The model may generate content that could be considered inappropriate in specific cultural contexts. Users should be aware of local cultural norms.

  3. Digital Divide : While this model helps bridge linguistic divides, access to the technology required to use this model may still be limited in some Bambara-speaking communities.

  4. Dialectal Variation : Bambara has several dialects, and this model may not perform equally well across all regional variations.

Additional Information

This model was developed as part of the MALIBA project, which aims to increase representation of African languages in AI systems. The complete implementation pipeline is available at https://github.com/sudoping01/instructions-gen .

The work is still in progress, with ongoing improvements planned for future releases.

Citation
@article{diallo2025bambara,
  title={Linguistically-Informed Large Language Models for Low-Resource Instruction Dataset Creation},
  author={Diallo, Seydou},
  journal={Unpublished manuscript},
  year={2025},
  month={July}
}
Framework Versions
  • PEFT 0.17.0
  • Transformers 4.55.2
  • PyTorch 2.6.0+cu124
  • Datasets 4.0.0
  • Tokenizers 0.21.4

Runs of sudoping01 maliba-llm on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About maliba-llm huggingface.co Model

More maliba-llm license Visit here:

https://choosealicense.com/licenses/gemma

maliba-llm huggingface.co

maliba-llm huggingface.co is an AI model on huggingface.co that provides maliba-llm's model effect (), which can be used instantly with this sudoping01 maliba-llm model. huggingface.co supports a free trial of the maliba-llm model, and also provides paid use of the maliba-llm. Support call maliba-llm model through api, including Node.js, Python, http.

sudoping01 maliba-llm online free

maliba-llm huggingface.co is an online trial and call api platform, which integrates maliba-llm's modeling effects, including api services, and provides a free online trial of maliba-llm, you can try maliba-llm online for free by clicking the link below.

sudoping01 maliba-llm online free url in huggingface.co:

https://huggingface.co/sudoping01/maliba-llm

maliba-llm install

maliba-llm is an open source model from GitHub that offers a free installation service, and any user can find maliba-llm on GitHub to install. At the same time, huggingface.co provides the effect of maliba-llm install, users can directly use maliba-llm installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

maliba-llm install url in huggingface.co:

https://huggingface.co/sudoping01/maliba-llm

Url of maliba-llm

Provider of maliba-llm huggingface.co

sudoping01
ORGANIZATIONS

Other API from sudoping01