The Latest AIs, every day
AIs with the most favorites on Toolify
AIs with the highest website traffic (monthly visits)
AI Tools by Apps
Discover the Discord of AI
AI Tools by browser extensions
GPTs from GPT Store
Discover The Best Model For AI
Top AI lists by month and monthly visits.
Top AI lists by category and monthly visits.
Top AI lists by region and monthly visits.
Top AI lists by source and monthly visits.
Top AI lists by revenue and real traffic.

The second generation of
Ąžuolas
, the Lithuanian ASR fine-tune series. v1
(
akisviete/azuolas-whisper-lt
)
was
whisper-large-v3
; v2 replaces the backbone with
Qwen/Qwen3-ASR-1.7B-hf
and
keeps everything that made v1 what it was: a full fine-tune on the complete
LIEPA-3
Lithuanian corpus — read, spontaneous and phonetic, ~
9,806 hours
.
Qwen3-ASR's technical report lists 52 supported languages, and
Lithuanian is
not among them
(
resolve_language()
raises on
lt
). v2 is therefore not a
refinement of an existing capability — it is a capability installed by
fine-tuning, and the training recipe reflects that (most importantly: a
learning rate 5× higher than the official SFT default, which is the single
largest measured lever — see
Training
).
The two generations split by job. Under a matched decode (beam 5, one version-matched stack) v1 is the more accurate model on the two large clean sets (8.93 vs 9.74 FLEURS, 5.09 vs 5.35 CV26). v2 rejects non-speech where v1 hallucinates over it , runs ~30 % faster under PyTorch at every decode setting tested, is 4.5× smaller, and is the integrated Subtitle Edit / CrispASR deployment path. One operational rule: run this model with beam search, not greedy — beam 5 is the decode every matched-decode comparison in this card uses and the one the CrispASR deployment path runs. The measured v2 vs v1 has the full picture.
Output is lowercase and unpunctuated
, following the LIEPA-3 label
convention, and uses the realised-speech convention — it writes
turim
, not
normative
turime
, because that is what the corpus labels do.
Evaluated on the author's
eval_compare.py
harness:
normalize_lt()
on both
sides,
jiwer.wer
/
jiwer.cer
. Both sets are
out of domain
: neither
FLEURS nor Common Voice is part of LIEPA-3, and neither shares its recording
conditions.
FLEURS LT — 985 clips, 2.96 h, 17,092 reference words
Common Voice 26.0 LT — 5,877 clips, 8.50 h, 39,653 reference words
(release
cv-corpus-26.0-2026-06-12
)
All numbers are beam 5 on one fresh version-matched stack (uv venv, torch 2.14.0+cu130, transformers 5.17.0, beam 5 wherever the architecture supports it):
| model (decode) | FLEURS WER / CER | CV26 WER / CER |
|---|---|---|
| Ąžuolas v1, large-v3 (beam 5) | 8.93 / 5.84 | 5.09 / 1.38 |
| Ąžuolas v2 1.7B — this model (beam 5) | 9.74 / 6.79 | 5.35 / 1.70 |
| kmynas-parakeet-lt-v4, TDT 0.6B (NeMo beam 5) | 10.56 / 6.18 | 6.56 / 1.42 |
| paprika v3, distilled (beam 5) | 11.60 / 6.41 | 8.08 / 1.72 |
Notes on the table.
Harness validation:
the same 2026-09-25 head-to-head
also ran both Ąžuolas generations on their canonical (greedy) decode, and
both reproduced their previously published figures to the hundredth (9.12 /
4.97 for v1, 9.74 / 5.11 for v2) — that is what makes the matched beam-5
comparison above trustworthy. Cross-harness deltas below ~0.5 points are
noise.
kmynas-parakeet-lt-v4
(a 3,087 h LIEPA-3 transducer fine-tune) is
greedy-only in the transformers path (
ParakeetForTDT
supports
GREEDY_SEARCH
only); its row above is the NeMo path (nemo 3.0.0), where
transducer beam 5 changes almost nothing versus greedy (10.56 / 6.60) —
for this TDT, beam search ≈ greedy — while the optimized NeMo decode
is the speed leader by far (on the 57-minute hard-chunk long-form run: 2.0 s
wall, 6,568 words, zero degeneration). It also emits punctuation, casing and
standardised endings;
normalize_lt
strips punctuation but not the
daryt→daryti
standardisation, so its WER carries a small convention tax.
paprika v3
is a distilled Whisper (32-layer large-v3 encoder + 4-layer
decoder, 1.6 GB) — the raw-speed option.
Reading: v1 leads the clean sets (−0.81 FLEURS / −0.26 CV26 vs next-best);
kmynas-v4 is the speed leader by far
; this model sits second on both
clean sets — ahead of kmynas-v4 (10.56 / 6.56) and paprika v3
(11.60 / 8.08) on both. Context from prior harnesses, not in
the table: stock
openai/whisper-large-v3
24.37 FLEURS / 30.07 CV26,
kmynas-v3 (v4's predecessor) 17.73 / 22.86.
The same ranking survives a completely different inference stack. Re-run through CrispASR (q8_0 GGUF, Vulkan, RTX 3080 Ti) over the same full sets:
| model (q8_0, CrispASR) | FLEURS WER / CER | CV26 WER / CER | speed |
|---|---|---|---|
| Ąžuolas v1 | 9.07 / 5.87 | 4.52 / 1.06 | 11.4× RT |
| Ąžuolas v2 1.7B (this model) | 9.79 / 6.24 | 5.03 / 1.28 | 13.7× RT |
q8_0 quantization is effectively lossless for this model: 9.79/5.03 (q8_0) against 9.74/5.11 (PyTorch bf16), inside the ~0.2 noise floor. The q8_0 GGUF (2.5 GB) is the artifact that runs in Subtitle Edit — see Deployment as GGUF .
Long-form — a 57-minute public event, unscripted. Vidas Mačiulis's book launch at Ąžuolyno biblioteka (57.1 min, audience, overlapping speech, applause). No human reference exists for this recording, so no WER is reported; what is measurable is coverage, degeneration and speed:
model (CrispASR,
--vad
)
|
sec | ×RT | words | coverage | distinct-5-gram | loop |
|---|---|---|---|---|---|---|
| Ąžuolas v2 1.7B (this model) | 292.3 | 11.7 | 6,596 | 99.8 % | 1.000 | none |
| Ąžuolas v1 | 299.6 | 11.4 | 6,631 | 99.8 % | 1.000 | none |
No degeneration, and the two models' outputs agree within a few percent. Qwen3-ASR's audio encoder handles long input at inference regardless of training clip length — a model trained almost entirely on ~5 s raw clips (an ablation, not a deployment candidate) ran the same 57-minute file with identical segment counts and zero looping.
Non-speech rejection — the axis where this model wins. Trained with a non-speech pool injected at a realised 1.40 %, it scored 100 % rejection, 0 characters on digital silence, faint hiss, room tone, ESC-50 and DEMAND (the last two never used in training).
Same-protocol head-to-head (20 s pure silence + 20 s white noise at −40 dB, PyTorch, both models):
| input | v1 (whisper) | v2 (this model) |
|---|---|---|
| pure silence | 22 chars | 0 chars |
| −40 dB noise | 414 chars of fluent hallucination | 0 chars |
v1 also invents up to 947 characters over faint hiss on its own card's 60 s protocol, and its no-speech gate does not fire — the production VAD exists to suppress exactly this. If your audio contains silence, noise or music, this is the model that stays quiet.
Lithuanian transcription of read and spontaneous speech: interviews, public events, parliamentary and broadcast audio, phone recordings. 16 kHz mono input (the processor resamples for you).
If you need punctuation, casing or speaker labels, those are separate models.
Requires
transformers >= 5.13.0
(Qwen3-ASR is natively supported from that
version):
pip install "transformers>=5.13.0"
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL = "akisviete/azuolas-qwen3-asr-1.7b-lt"
processor = AutoProcessor.from_pretrained(MODEL)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL, dtype=torch.bfloat16, device_map="cuda")
inputs = processor.apply_transcription_request(
audio="clip.wav",
language="Lithuanian", # force the language; auto-detect also works
).to(model.device, model.dtype)
out_ids = model.generate(**inputs, max_new_tokens=256, num_beams=5)
gen_ids = out_ids[:, inputs["input_ids"].shape[1]:]
# raw output is language Lithuanian<asr_text>…
print(processor.decode(gen_ids, return_format="transcription_only")[0])
return_format="parsed"
gives a
{"language": …, "transcription": …}
dict
if you also want the detected language.
Use beam search, not greedy.
This model is an encoder-LLM and is
sensitive to the decode setting in a way Whisper models are not — on
in-distribution sets the difference is small, but far out of distribution
the LLM decoder can collapse on rare words and proper names under greedy.
num_beams=5
is the setting every matched-decode comparison in this card
uses and the one the CrispASR deployment path runs.
Decode in windows of
at most ~22–26 s
, cut at the midpoints of pauses
found by
Silero VAD
, or use an engine that does this for you (CrispASR
with
--vad -vm silero
). The training windows average 25.44 s, so this is
the regime the model was built for.
Fixed-stride chunking is measured to under-transcribe: on 180 s of Lithuanian speech, 30 s/28 s-stride chunks reached 82.4 % agreement with the reference against 94.9 % for pause-cut ~22 s windows — it does not truncate, it silently omits runs of 10–16 words from the middle.
A q8_0 GGUF of these weights (2.5 GB) is the deployed artifact for Subtitle
Edit's
Crisp ASR Qwen3
engine (CrispASR's
convert-qwen3-asr-to-gguf.py
,
llama.cpp-style tensor names). Verified end to end: 180 s of audio in 49 s
(3.7× RT, q8_0, Vulkan, RTX 3080 Ti), word-level timestamps via
Qwen3-ForcedAligner, diacritics intact. Always pass
--vad
(CrispASR's
qwen3 path runs beam search + VAD, which is the production decode this model
wants — see
Results
). A GGUF of this
exact checkpoint can be published alongside the HF weights if requested.
Measured reference points:
| path | weights in memory | what was measured |
|---|---|---|
| PyTorch bf16, beam 5, clips | ≈3.4 GB (1.7 B params × 2 bytes) | the beam-5 evaluation ran at batch 12 with ~19 GB free on a 96 GB card time-sliced with a 78 GB vLLM tenant; greedy batch 32 / beam batch 16 OOM'd in the same headroom |
| PyTorch bf16, whole file | ≈3.4 GB + audio-feature prefill | prefill of 57 min of audio OOM'd with ~19 GB free — long audio must be VAD-chunked (see above) |
| GGUF q8_0 (CrispASR) | 2.5 GB | runs on a 12 GB RTX 3080 Ti with the forced aligner resident; the f16 GGUF (4.4–4.7 GB) does not fit 12 GB once KV cache + aligner (~5 GB overhead) are added |
| Training (for reference) | bf16 | peak 77.93 GB on one 96 GB GPU |
Rule of thumb derived from those points: the GGUF path wants ~8 GB minimum / 12 GB comfortable ; the PyTorch path wants 16 GB for comfortable beam-5 clip inference. The minimum figures are derived, not measured — everything in the table is measured.
num_beams=5
is the production decode — this card's numbers and the CrispASR path.
dial
subset (~100 h) was excluded: 99.95 %
of its transcripts carry phonetic dialect annotation rather than standard
orthography.
Full SFT of
Qwen/Qwen3-ASR-1.7B-hf
(no LoRA — see below), 1 epoch on the
full 9,806 h corpus, ~16 h wall time on one 96 GB GPU.
| base |
Qwen/Qwen3-ASR-1.7B-hf
(full fine-tune, all weights)
|
| data | LIEPA-3, 9,806 h — read 5,230 h (53.3 %) + spontaneous 4,359 h (44.5 %) + phonetic 217 h (2.2 %), Whisper-prep fused ~30 s windows (mean 25.44 s), streamed at the corpus's natural subset proportions |
| non-speech pool |
19,755 clips / 42.57 h (synthetic, MUSAN, AudioSet — all CC BY 4.0 or generated), injected at
nonspeech_rate 0.0142
,
realised 1.40 %
— this is what buys the 100 % rejection
|
| optimiser |
AdamW, β (0.9,
0.98
), ε
1e-9
, weight decay
0.1
,
max_grad_norm 1.0
|
| lr | 1e-4 , linear decay to zero, warmup ratio 0.02 (~220 steps) |
| batch | 16 × 8 accumulation = effective 128 |
| steps | 10,995 |
| precision | bf16, FlashAttention 2, no gradient checkpointing |
| checkpoint | bf16 |
The official Qwen3-ASR SFT recipe specifies lr 2e-5 . It is ~5× too low for this task:
| lr | FLEURS | CV26 |
|---|---|---|
| 2e-5 (official) | 12.95 | 9.83 |
| 1e-4 (this model) | 9.74 | 5.11 |
| 2e-4 | 10.97 | 5.96 |
The 2e-5 convention assumes the model already has the capability and you are
nudging it. Lithuanian is absent from the base model's 52 languages, so the
weights must travel much further than a domain nudge can in ~1,100 steps —
and
max_grad_norm 1.0
was clipping the under-warmed 2e-5 run most of the
time, compounding the undershoot. The curve peaks at 1e-4; 2e-4 is a real
regression (+1.23 FLEURS), so this is a peak, not a plateau.
Same corpus, same series, different backbone — run side by side on one PyTorch harness on 2026-09-25, including the full five-model beam-5 leaderboard ( Results ) and a worst-case long-form stress test (hard 30 s chunks, greedy).
On accuracy, v1 is the better ASR model on the two large clean sets , under a matched decode (beam 5, one fresh version-matched stack for both): 8.93 vs 9.74 FLEURS, 5.09 vs 5.35 CV26.
v2's defensible advantages are the other four axes, all measured:
prompt
),
DPO-tunable, usable for downstream reasoning over the transcript. v1 is a
seq2seq model that does exactly one job.
Rule of thumb. For verbatim reference transcription of clean Lithuanian speech — e.g. the TTS data pipeline — v1 keeps the edge on the big clean sets. For the interactive / Subtitle Edit / CrispASR deployment, where latency, silence robustness and footprint matter, v2 is the right choice — run it with beam search, not greedy.
Against the field, on the matched beam-5 leaderboard: second on both clean sets — behind v1 (8.93 / 5.09 vs 9.74 / 5.35) — and ahead of kmynas-v4 (10.56 / 6.56) and paprika v3 (11.60 / 8.08) on both. It is the only one of the field that ships a working non-speech guarantee, and the only encoder-LLM.
Base model
Qwen/Qwen3-ASR-1.7B-hf
—
https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf
(Apache-2.0). Qwen3-ASR technical report;
Qwen3-ForcedAligner-0.6B
for
word timestamps.
Models compared
akisviete/azuolas-whisper-lt
(Ąžuolas v1) —
https://huggingface.co/akisviete/azuolas-whisper-lt
paprika-whisper-lt-v3
—
https://huggingface.co/kristijonas/paprika-whisper-lt-v3
kmynas-parakeet-lt-v3
—
https://huggingface.co/kristijonas/kmynas-parakeet-lt-v3
kmynas-parakeet-lt-v4
—
https://huggingface.co/kristijonas/kmynas-parakeet-lt-v4
Data
www.raštija.lt
).
Copy used:
https://huggingface.co/datasets/meldynamics/liepa-3
cv-corpus-26.0-2026-06-12
—
https://commonvoice.mozilla.org/en/datasets
Tooling
qwen3
/
qwen3-1.7b
backends used for the deployment numbers)
Weights
CC BY 4.0
, inheriting the LIEPA-3 corpus attribution chain
(Vilnius University / raštija.lt). Base model
Qwen/Qwen3-ASR-1.7B-hf
is
Apache-2.0. The non-speech pool is all CC BY 4.0 or generated.
Questions and suggestions: [email protected] .
azuolas-qwen-lt huggingface.co is an AI model on huggingface.co that provides azuolas-qwen-lt's model effect (), which can be used instantly with this akisviete azuolas-qwen-lt model. huggingface.co supports a free trial of the azuolas-qwen-lt model, and also provides paid use of the azuolas-qwen-lt. Support call azuolas-qwen-lt model through api, including Node.js, Python, http.
azuolas-qwen-lt huggingface.co is an online trial and call api platform, which integrates azuolas-qwen-lt's modeling effects, including api services, and provides a free online trial of azuolas-qwen-lt, you can try azuolas-qwen-lt online for free by clicking the link below.
azuolas-qwen-lt is an open source model from GitHub that offers a free installation service, and any user can find azuolas-qwen-lt on GitHub to install. At the same time, huggingface.co provides the effect of azuolas-qwen-lt install, users can directly use azuolas-qwen-lt installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
