Private / internal model.
Evaluate on your own data before production use.
Results (932-clip held-out test set)
WER
CER
Base (cohere-transcribe-arabic, zero-shot)
0.457
0.174
This model (full fine-tune)
0.357
0.137
~22% relative WER cut over the (already strong) Arabic-specialized base. This is
the model the base
should
be fine-tuned into: an earlier
LoRA
attempt on the
same data made it
worse
(0.457 → 0.510, overfit). A
full
fine-tune with a
low LR and best-checkpoint selection was the fix.
⚠️
This model overfits easily.
Eval loss bottomed at
step 1000
(0.342) and
rose afterward, while WER stayed flat — the saved weights are that best checkpoint
(
load_best_model_at_end
). Watch
WER
, not just loss, if you train further.
Model comparison — all Arabic ASR models
Same 932-clip held-out test set, same
clean_text
scoring (strip tashkil + tags,
keep punctuation + dialect spelling) — so every row is directly comparable.
Model
Params
Zero-shot WER
Fine-tuned WER
CER (best)
whisper-large-v3-turbo
🏆
809M
0.590
0.344
0.115
cohere-transcribe-arabic
2.0B
0.457
0.357
0.137
whisper-medium
769M
0.717
0.358
0.123
nemotron-3.5-asr (streaming)
638M
0.592
0.422
—
whisper-small
244M
~0.77
0.428
0.151
qwen3-asr-0.6b
938M
0.756
0.676
0.408
qwen3-asr-1.7b
1.7B
—
training
—
Best fine-tuned:
whisper-large-v3-turbo
(WER 0.344), with
cohere-transcribe-arabic
a close second (
0.357
).
Best zero-shot:
cohere-transcribe-arabic
(0.457, Arabic-specialized). A
full
fine-tune (all ~2B params, low LR) now
improves
it to
0.357
; an earlier 32 GB
LoRA
attempt had instead
degraded
it (0.510, overfit) — full-parameter tuning
with best-checkpoint selection was the fix.
Streaming / low-latency:
nemotron-3.5-asr
.
Parakeet-TDT-0.6b-v3 was tried but abandoned (European-only pretraining;
cross-lingual transfer to Arabic converged far too slowly, WER ~0.93 after 3k steps).
Augmented
: each clean clip is expanded with variants (noise, music, speed,
voice, reverb/codec) via an
augmentation
column. Derived from
oddadmix/dialectal-arabic-lahgtna-v2-smaller
.
Kept
dialectal consonants (گ ڨ چ پ ژ) — they encode real phonemes — and
sentence punctuation. Output is
undiacritized
.
Filtered
: dropped clips > 30 s and a few corrupt rows (huge-text blobs).
~62% of rows had harakat, ~28% had non-verbal tags before cleaning.
→
36,769 train / 932 test
after filtering.
⚠️
Leakage note:
the set is augmentation-expanded from shared source clips,
so train/test may share source audio → held-out numbers can be optimistic. Use a
leakage-free split for true numbers.
Training
Base
:
CohereLabs/cohere-transcribe-arabic-07-2026
·
full fine-tune
(no LoRA)
Hardware
: a single
48 GB RTX A6000
(~43 GB peak, ~8.3 s/step)
Metric
: WER/CER on cleaned references (clean-text normalized), computed by
per-sample
generation (batched generation garbles this model's output)
Fine-tuning (reproduce)
The exact fine-tuning code is bundled in this repo (
train_cohere_full.py
,
normalize.py
,
eval_hf_asr.py
,
ds_zero2.json
,
setup.sh
) plus
requirements.txt
.
See
FINETUNE.md
for the full walkthrough + lessons learned.
Trained on
oddadmix/dialectal-arabic-lahgtna-v2-smaller-augmented
(
private
) —
swap in any HF audio dataset with
audio
+
text
columns.
normalize.py
is the
shared text cleaning (strip tashkil + non-verbal tags, keep dialectal letters گ ڨ چ).
For less memory / multi-GPU,
accelerate launch train_cohere_full.py --deepspeed ds_zero2.json ...
.
Usage
# needs transformers from source (ships the cohere_asr architecture)import torch, torchaudio
from transformers import AutoProcessor, CohereAsrForConditionalGeneration
repo = "oddadmix/cohere-transcribe-arabic-07-2026-dialectal"
proc = AutoProcessor.from_pretrained(repo)
model = CohereAsrForConditionalGeneration.from_pretrained(
repo, torch_dtype=torch.bfloat16).to("cuda").eval()
wav, sr = torchaudio.load("clip.wav") # 16 kHz mono
inp = proc(wav.mean(0).numpy(), sampling_rate=16000,
language="ar", return_tensors="pt").to(model.device)
inp["input_features"] = inp["input_features"].to(model.dtype) # keep ids long
ids = model.generate(**inp, max_new_tokens=256)
print(proc.tokenizer.batch_decode(ids, skip_special_tokens=True)[0])
Decode
one bare array at a time
. Passing a list +
padding=True
triggers a
batched/padded path that badly degrades this model's generation (WER 0.45 → 0.65+).
Learnings & notes
A strong specialist can be
hurt
by fine-tuning.
Cohere's Arabic model has the
best
zero-shot
WER of everything tried (0.457). LoRA on this data overfit and made
it worse (0.510); only a
full
fine-tune with a
low LR (5e-6)
and
best-checkpoint selection
turned it into a net win (0.357).
Watch WER, not loss.
eval_loss
bottomed at step 1000 and rose while WER held —
more steps would have overfit.
load_best_model_at_end
keeps the right weights.
Data cleaning matters most
: keeping dialectal letters + stripping tashkil/tags
(rather than aggressive letter-folding) was the biggest lever for a multi-dialect model.
The processor concatenates the language prompt ⊕ transcript for the model's causal-LM
loss — handled in the collator, no action needed.
Limitations
Trained on augmented, partly synthetic multi-dialect data; real-world dialect
coverage varies (Maghrebi is the hardest).
Output is
undiacritized
and lower-cased for Latin tokens.
Offline model (≤30 s chunks); not streaming.
Overfits easily on smaller/noisier data — use a low LR and keep the best checkpoint.
Runs of oddadmix cohere-transcribe-arabic-07-2026-dialectal on huggingface.co
331
Total runs
6
24-hour runs
22
3-day runs
28
7-day runs
193
30-day runs
More Information About cohere-transcribe-arabic-07-2026-dialectal huggingface.co Model
More cohere-transcribe-arabic-07-2026-dialectal license Visit here:
cohere-transcribe-arabic-07-2026-dialectal huggingface.co is an AI model on huggingface.co that provides cohere-transcribe-arabic-07-2026-dialectal's model effect (), which can be used instantly with this oddadmix cohere-transcribe-arabic-07-2026-dialectal model. huggingface.co supports a free trial of the cohere-transcribe-arabic-07-2026-dialectal model, and also provides paid use of the cohere-transcribe-arabic-07-2026-dialectal. Support call cohere-transcribe-arabic-07-2026-dialectal model through api, including Node.js, Python, http.
cohere-transcribe-arabic-07-2026-dialectal huggingface.co is an online trial and call api platform, which integrates cohere-transcribe-arabic-07-2026-dialectal's modeling effects, including api services, and provides a free online trial of cohere-transcribe-arabic-07-2026-dialectal, you can try cohere-transcribe-arabic-07-2026-dialectal online for free by clicking the link below.
oddadmix cohere-transcribe-arabic-07-2026-dialectal online free url in huggingface.co:
cohere-transcribe-arabic-07-2026-dialectal is an open source model from GitHub that offers a free installation service, and any user can find cohere-transcribe-arabic-07-2026-dialectal on GitHub to install. At the same time, huggingface.co provides the effect of cohere-transcribe-arabic-07-2026-dialectal install, users can directly use cohere-transcribe-arabic-07-2026-dialectal installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
cohere-transcribe-arabic-07-2026-dialectal install url in huggingface.co: